# Designing browser music controls around the next beat

> Source: <https://dev.to/paulgendek/designing-browser-music-controls-around-the-next-beat-29f6>
> Published: 2026-10-01 21:22:57+00:00

When I am playing guitar alone, changing a backing groove usually means moving a hand from the instrument to the laptop. A button can be easy to click and still interrupt the activity it is supposed to support. That problem led me to Mock Band, a browser rhythm section, and to a useful distinction: accepting a command, applying a musical change and making that change understandable to the player are separate jobs.

I supplied the concept, product direction and requirements, tested physical hardware and retained merge approval. AI tools implemented changes and ran automated checks. This article is an AI-drafted adaptation of my original project account and the current case study. The implementation details below describe the system those tools helped implement.

The initial idea, recorded in a note in September 2023, was to choose drums, bass and guitar accompaniment, repeat a groove and then reshape it through instructions. A player could change key or move to half-time while keeping both hands on the guitar. A footswitch would activate voice listening.

The player remained responsible for the riff and musical direction. The product did not need to listen to the guitar and infer its performance to serve that goal. It needed to accept an intentional change with less interruption.

That distinction reduced the interaction problem to something concrete. How does a person indicate a change? Which parts of that change can happen now? Which should wait for a musical boundary? What should remain usable when voice or a controller is unavailable?

I chose the browser after researching WebMIDI and browser music APIs. Voice recognition, MIDI support and permissions still vary across browser and device combinations. Visible controls and text therefore remain useful ways to operate the band.

The interface and application state use React and TypeScript. Text and recognized speech share a command vocabulary. A typed tempo instruction and its recognized spoken equivalent can reach the same musical action without maintaining two separate definitions of what the instruction means.

MIDI controller input has a different shape. A mapped controller action can invoke a scene or transport handler directly. It does not need to become a sentence and pass through a text parser first.

The useful point of convergence is shared application state and provider handlers. Those handlers connect transport, musical changes and scene events to the rest of the system. Sharing that destination gives the input methods a common effect while preserving the differences in how the request arrived.

This matters when extending a product. Adding an input method should not require inventing another version of scene recall or tempo changes. It also should not require squeezing every physical gesture into a text-command metaphor.

The current audio bridge passes changes to an AudioWorklet. The worklet renders drums, bass and rhythm accompaniment and handles their timing away from interface rendering. Its live schedule can also feed WebMIDI output through the bridge.

[MDN's AudioWorklet reference](https://developer.mozilla.org/en-US/docs/Web/API/AudioWorklet) explains the separate Web Audio execution context. That separation helps explain this architecture, but the use of a worklet alone does not establish that a particular application has no timing problems.

The interface has to respond to a request. The audio engine has to decide when an allowed change takes effect. In this system, a playing change can queue for a beat or bar boundary; a stopped change can apply immediately.

Consider a tempo adjustment requested halfway through a bar. The player has made a valid request, but applying it at an arbitrary instant may disrupt the phrase. Queuing it gives the music a meaningful point at which to change. This is a design choice about musical behavior, with implementation consequences for state.

A delayed musical change creates more than one truthful description of the current situation.

| State | What it means | What the player needs to understand | 
|---|---|---|
| Selected | The player has requested a new value. | The intended next setting. | 
| Pending | The change is waiting for a musical boundary. | Why the current sound may still reflect the previous value. | 
| Applied | The audio engine has adopted the change. | Which setting is active in the music. | 

These are design distinctions, not three competing copies of the truth. A selected value can be correct as an intention while the old value remains correct as the applied sound.

For a hypothetical example, imagine changing from one tempo to another during playback with a bar-boundary change pending. If the interface presents the new value as already active, a player may think the engine ignored it. If it hides the request until the boundary, the player may repeat it. The product needs to communicate the relationship between intention and execution.

The lesson extends beyond music. Whenever an operation is accepted now but takes effect later, the interface needs to describe both events honestly. In music, the delay has an audible consequence, so the distinction becomes difficult to ignore.

During early development, I tested a purchased Bluetooth footswitch and found that its behavior was a toggle. The interaction had assumed a hold-to-talk gesture.

That was not a problem I could resolve by looking only at a button handler. The intended gesture and the actual device disagreed. I fed the finding back into Cursor and refined the interaction around a tap that starts listening while recognition is awaited. Getting it usable required debugging and repeated manual testing.

Scene recall with the footswitch was a satisfying early result for me. That is a historical acceptance observation, not certification of the current version or every controller.

The important requirement came from trying the hardware in the activity. With both hands on a guitar, a switch has to support a gesture the player can actually perform and recognize. A software abstraction can describe a press cleanly while missing how the physical control behaves.

The current product provides ten local scenes, with JSON export and import for backup or moving configurations between browsers. Keeping musical state in the browser avoids an account and cloud setup step. It also makes the storage boundary part of the experience: a scene saved in one browser is not automatically present in another.

Local state should not be confused with offline support for every feature. Browser speech recognition has its own capabilities and dependencies. A local scene system does not make a browser's recognition service work offline.

Validation has similar boundaries. Agents handled automated checks. I supplied manual acceptance testing and merge decisions. Those activities can support different claims, and none should stand in for all the others.

I reported a successful live MIDI output test through an IAC bus into Logic Pro on September 29, 2026. That is useful evidence for that route. It does not establish behavior with every physical receiver or controller. An interface screenshot can establish a visible control or layout, but cannot establish that the music sounds right.

The next useful checks are therefore about the actual interaction: can a player request a change without losing the phrase, understand when it takes effect and recover when an optional input method is unavailable? The current version has not yet been part of my regular practice beyond QA. Those practice questions remain work to do.

Building a browser music tool made the input and timing boundary unusually tangible for me. The system must accept the player's intention, preserve a musically sensible transition and communicate the difference. Real hardware then tests whether that intention was represented correctly in the first place.
