Skip to content

oto: add Player.UnplayedSize, reported on Android - #307

Open
marrasen wants to merge 5 commits into
ebitengine:mainfrom
marrasen:android-output-latency
Open

marrasen wants to merge 5 commits into
ebitengine:mainfrom
marrasen:android-output-latency

Conversation

@marrasen

@marrasen marrasen commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

This change and its description were written by Claude, an AI agent, working with @marrasen on gunim, a GUI framework that plays its sound through oto. Marcus tested it on his phone.

What issue is this addressing?

Updates #311

What type of issue is this addressing?

feature

What this PR does | solves

An application that draws the sound it plays needs to know which part of a player's data is being heard now. BufferedSize counts only what the player still holds, but after the mux the sound takes a while longer to reach the listener: about 250 ms over Bluetooth earbuds. This adds one method:

// UnplayedSize returns the byte size of the data read from the source that is not heard yet: the buffered data, and
// the data sent to the audio hardware that it has not played yet, also after Pause and after the end of the source.
// Where the platform does not report how long the audio hardware takes to play, UnplayedSize returns the same as
// BufferedSize. Seek and Reset forget the data already sent.
func (p *Player) UnplayedSize() int

So the part of the source being heard now is the bytes read from it less UnplayedSize, after Pause, right after Play, and at the end of the source alike.

The mux

  • Mux counts the samples ReadFloat32s mixes.
  • Each player records the spans of the mix its data went into, while they may be unheard. Spans the player fills back to back merge, so a player that plays on keeps one.
  • A driver reports, through Mux.SetDelayFunc, how many of the frames the mux has mixed are not heard yet. Where it reports nothing, UnplayedSize returns what BufferedSize does.
  • The buffer and the spans are read under the player's lock together, so a read in between cannot move data from one to the other.
  • Seek and Reset forget the spans, as their data is from the old position.

Android

The binding reports the delay. Every 100 ms, LoopRead measures when the frame written to the fifo next will be heard, from calculateLatencyMillis and the frames queued, and Delay counts on from that pair, so the fifo and the stream are sampled together.

  • The stream is asked outside mutex_. Its pointer and a generation counter are copied under the lock, and the result is kept only if the generation is unchanged and no callback ran meanwhile.
  • The measurement is forgotten wherever the stream is dropped, paused or replaced.
  • A negative latency, as around an underrun, is clamped to zero.
  • OpenSL ES has no timestamps, so there the frames the stream holds are the estimate.

The other drivers report nothing yet. Each has a TODO beside its mux.New, naming what it would use. Context.OutputLatency, from the first version of this PR, is gone.

Testing

internal/mux/unplayed_test.go drives the mux as a driver does, with a delay it controls: no delay reported, steady play, right after Play, after Pause, at the end of the source, and after Seek. The tests run in testing/synctest bubbles and wait for the mux's loop with synctest.Wait, so their results depend on no machine's scheduler. They pass with -race -cpu=1,4.

On a Pixel 8 Pro (Android 16) with Sony WF-1000XM4 earbuds, in a music player built on gunim, Marcus checked the spectrum and the animations against the sound: in step during play, after pausing and resuming (from the lock screen too), across a change of track, and with the earbuds taken out and put back.

On the Android emulator, a stress app played through gunim and this change for 15 minutes, 2,300 random actions: seeks, mostly near the end of a track, the next track queued, pauses, the speaker suspended and resumed, its buffer resized, bursts of garbage, and the screen and the app's foreground changed from outside. What it computed as heard never fell further behind than the buffer and the device's delay, 610 ms, and no measurement failed.

The branch is based on main, and the workflow passed on it in my fork, on all nine jobs: https://github.com/marrasen/oto/actions/runs/38035884350

🤖 Generated with Claude Code

@marrasen marrasen changed the title Claude bug report: no way to know how late the sound is heard, so visuals run ahead over Bluetooth oto: add Context.OutputLatency, reported on Android Oct 7, 2026
@marrasen
marrasen force-pushed the android-output-latency branch from c43e32d to cf783d0 Compare October 7, 2026 05:55
@hajimehoshi

Copy link
Copy Markdown
Member

Resolve the conflicts.

Marcus Johansson and others added 2 commits October 8, 2026 08:05
OutputLatency returns how long the sound the context has read from its
players takes to be heard: what the fifo queues, and what the stream
holds, as its timestamp says. On Android the timestamp counts a
Bluetooth headset's own delay where the headset reports it. LoopRead
measures it every 100 ms. Other platforms report nothing yet.

An application that shows the sound it plays, as a spectrum or a beat,
needs it to show what is heard rather than what was read: over
Bluetooth earbuds the difference is a quarter of a second.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
OutputLatency is declared in context.go and calls the driver's
context, as Suspend, Resume and Err do. Android reports the latency,
and the other drivers report none for now; each is where a later
change adds its platform. outputlatency_android.go and
outputlatency_other.go are gone.

TestOutputLatency asks the test's context: a latency reported is zero
or more, and one not reported is zero.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@marrasen
marrasen force-pushed the android-output-latency branch from cf783d0 to e478e38 Compare October 8, 2026 06:06
@marrasen

marrasen commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Rebased onto 91d77a3, now that #306 is in. The conflicts were all in binding_android.cpp, where both changes add code at the same spots: buffers_ready_ is now set at the end of PrepareBuffersLocked, since the read thread starts in EnsureStreamLocked; a reopened stream's latency is reset after ConfigureRefillLocked; and LoopRead measures the latency before its refill logic. The workflow passed on it in my fork, on all nine jobs: https://github.com/marrasen/oto/actions/runs/37735789308

Reply written by Claude (Anthropic), an AI agent, and posted at Marcus's request.

@hajimehoshi

Copy link
Copy Markdown
Member

Cannot we implement the same things for other platforms? It's ok to focus on Android in this PR, but at least we should leave TODO comments

Each driver that reports no latency yet has a TODO naming what it
would use: PulseAudio's latency query and snd_pcm_delay on Linux,
IAudioClock and waveOutGetPosition on Windows, the audio queue's time
and the device's latency on macOS and iOS, and the AudioContext's
latencies on the web.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@hajimehoshi hajimehoshi left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from the shape of the API, I found some issues in the Android measurement. Most of this code would carry over to the player-level design, so they apply either way.

This review was drafted with Claude (Claude Code).

Comment thread internal/oboe/binding_android.cpp Outdated
// stream_latency_frames_ is how many frames the stream playing holds that are
// still to be heard, as its timestamp last said, or -1 while it has not said.
// LoopRead measures it every kLatencyEvery.
std::atomic<bool> buffers_ready_{false};

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not needed. Only LoopRead stores a non-negative stream_latency_frames_, and it starts after fifo_ is made. If Latency checks stream_latency_frames_ first, a value of 0 or more already means fifo_ is ready.

Comment thread internal/oboe/binding_android.cpp Outdated
}
ConfigureRefillLocked();
// The stream is new, and has not said how long it takes yet.
stream_latency_frames_.store(-1);

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the only place the value is reset. After onErrorAfterClose, a failed restart, or Pause, the last value is still reported as valid, plus a fifo that has filled up because nothing drains it. Could it be reset wherever stream_ is dropped or leaves kRunning?

Comment thread internal/oboe/binding_android.cpp Outdated
if (stream < 0) {
return -1;
}
return stream + fifo_->getFullFramesAvailable();

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

stream was measured up to 100 ms ago, and the fifo count is from now. Each jumps by a burst at every callback, in opposite directions (Oboe's documentation says an output stream's latency "will increase abruptly when you write data to it"), so the sum is smooth only when both are sampled together. Otherwise it can be off by up to a burst, which is tens of ms over Bluetooth. How about recording, at each measurement, the fifo's write counter and when that frame will be heard, and computing the latency from that pair here?

Comment thread internal/oboe/binding_android.cpp Outdated
measured = now;
// The stream is reached under mutex_, which Pause and Resume hold only
// briefly; a measurement is skipped rather than waited for.
std::unique_lock<std::mutex> lock{mutex_, std::try_to_lock};

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This holds mutex_ across calculateLatencyMillis, which queries AAudio. The comment in StartLocked says nothing under mutex_ may wait on a device, as Pause and Resume take it on the UI thread. try_lock keeps this thread from waiting for others, but not others from waiting for it, and the fifo is not refilled meanwhile either. How about copying stream_ and a generation counter under the lock, querying the copy after releasing it, and storing the result only if the generation is unchanged?

Comment thread internal/oboe/binding_android.cpp Outdated
// briefly; a measurement is skipped rather than waited for.
std::unique_lock<std::mutex> lock{mutex_, std::try_to_lock};
if (lock.owns_lock() && stream_ && state_ == State::kRunning) {
if (auto ms = stream_->calculateLatencyMillis(); ms) {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oboe implements calculateLatencyMillis only for AAudio. OpenSL ES returns ErrorUnimplemented, and has no getTimestamp either. AudioApiForSdk uses OpenSL ES below Android 11, and StartOrDeferLocked falls back to it when AAudio refuses the configuration, so nothing is ever reported on those devices. Could we estimate it there, e.g. from the stream's buffer size plus the fifo? Otherwise the documentation should say so.

Comment thread internal/oboe/binding_android.cpp Outdated
std::unique_lock<std::mutex> lock{mutex_, std::try_to_lock};
if (lock.owns_lock() && stream_ && state_ == State::kRunning) {
if (auto ms = stream_->calculateLatencyMillis(); ms) {
stream_latency_frames_.store(

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

calculateLatencyMillis is the time the next frame will be heard minus now, with no clamping, so it can be negative, e.g. around an underrun. Latency treats any negative value as not reported yet, so the result flips to unavailable for at least 100 ms. Clamping it to 0 here would avoid that.

@hajimehoshi

Copy link
Copy Markdown
Member

Thank you for working on this. Before going into the details, I'd like to settle the shape of the API.

Context or player

The delay after the mux is shared by all the players, so measuring it per context in each driver makes sense. But apps need it per player: which part of this player's data is being heard now. Ebitengine's audio.Player.Position needs that, and so do gunim's visuals. Subtracting Player.BufferedSize and Context.OutputLatency from the bytes read gives the right answer only while the player keeps playing:

  • After Pause, the mux stops mixing the player at once, but what is already queued keeps playing for one latency. The computed position stops one latency short of what was heard, and stays off until one latency after the player resumes. Over Bluetooth that is about 250 ms.
  • Right after Play, the computed position is negative for one latency.
  • At the end of the source, the player is paused as soon as its last bytes are mixed, so IsPlaying becomes false about 250 ms before the sound ends.
  • BufferedSize and OutputLatency are read at different times, and a 10 ms read can move data from one to the other between the two calls.

Only the mux can get these right, so how about this?

  • Internally, each driver reports how many of the frames the mux has mixed are not heard yet. Most of the Android code in this PR carries over.
  • The mux counts the frames it mixes, and each player keeps track of which mixed frames its bytes went into, for as long as they can still be in flight.
  • A new Player method returns the bytes read from the source that are not heard yet. Just a suggestion, but it could be UnplayedSize() int. Where the driver reports nothing, it returns the same as BufferedSize, so no bool is needed.

Then Ebitengine only has to call the new method instead of BufferedSize to compute the position.

I'd rather not make Context.OutputLatency public for now. If a use apart from players comes up, it can be added later, but it cannot be removed in v3.

Unit

Bytes, like BufferedSize and SetBufferSize. Callers subtract it from the bytes read from the source, and bytes stay exact whole frames. Between the drivers and the mux, frames rather than time.Duration, which also avoids converting ms to frames to Duration and back.

Name

With this, OutputLatency goes away. Even for a context-level value I'd avoid it: Web Audio's AudioContext.outputLatency and iOS's AVAudioSession.outputLatency leave out the app's own buffering, and this value includes it, so the same name would mean something different.

Just a suggestion, nothing decided: if a context-level value is made public, OutputDelay might fit. ALSA's snd_pcm_delay means the same thing: how long a frame written now takes to be heard.

This comment was drafted with Claude (Claude Code).

…tency

The delay after the mux is shared by the players, but an app needs it
per player: which of this player's data is heard now, also after Pause,
right after Play, and at the end of the source.

- The mux counts the samples it mixes, and each player records the
  spans of the mix its data went into, while they may be unheard.
- Each driver reports, through Mux.SetDelayFunc, how many of the frames
  the mux has mixed are not heard yet. Android does; the others have a
  TODO naming what they would use.
- Player.UnplayedSize returns the bytes read from the source that are
  not heard yet: what BufferedSize returns, and what was sent on and
  not played. Where the driver reports nothing, it returns what
  BufferedSize does. Seek and Reset forget what was sent.
- Context.OutputLatency is gone.

On Android, LoopRead measures every 100 ms when the frame written to
the fifo next will be heard, and Delay counts on from that pair, so the
fifo and the stream are sampled together. The stream is asked outside
mutex_, and a measurement is kept only if the stream is unchanged and
no callback ran meanwhile. It is forgotten wherever the stream is
dropped, paused or replaced, clamped at zero, and estimated from the
stream's buffer on OpenSL ES, which has no timestamps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@marrasen marrasen changed the title oto: add Context.OutputLatency, reported on Android oto: add Player.UnplayedSize, reported on Android Oct 9, 2026
The UnplayedSize tests run in testing/synctest bubbles and wait for the
mux's loop with synctest.Wait, so their results depend on no machine's
scheduler. Their cleanup lets the loop see Stop after the moment it
sleeps once a source has ended.

sentSpan's literal has a field a line. Doc comments say what each
thing does: how spans merge is said where they merge, and
ForgetDelayLocked's says only what it forgets. MeasureDelay says why a
few tries are enough.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@marrasen

marrasen commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks, this is a better shape. 3d1f768 reworks the PR along your outline, and 8913cd7 runs its tests in virtual time:

  • The drivers report, through an internal Mux.SetDelayFunc, how many of the frames the mux has mixed are not heard yet. Only Android does so far; each other driver has a TODO beside its mux.New.
  • The mux counts the samples it mixes, and each player records the spans of the mix its data went into, while they may be unheard.
  • Player.UnplayedSize() int returns the bytes read from the source that are not heard yet, and what BufferedSize returns where the driver reports nothing. The buffer and the spans are read together under the player's lock. One choice of ours: Seek and Reset forget the spans, as their data is from the old position.
  • Context.OutputLatency is gone, and between the driver and the mux everything is in frames.

The six points from your review of the Android code are all in: no buffers_ready_; the measurement is forgotten wherever the stream is dropped, paused or replaced; the fifo and the stream are sampled together, as a pair of the write counter and when that frame will be heard; the stream is asked outside mutex_, with a generation check; OpenSL ES estimates from the stream's buffer; and a negative latency is clamped to zero.

The new tests in internal/mux cover no delay reported, steady play, right after Play, after Pause, the end of the source, and Seek, in testing/synctest bubbles. Marcus checked it on his phone with Bluetooth earbuds, in a music player built on gunim: in step during play, after pause and resume (from the lock screen too), across a track change, and with the earbuds out and back in. On the emulator, a stress app ran 2,300 random seeks, pauses and suspends over 15 minutes without falling behind. The workflow passed on all nine jobs in my fork: https://github.com/marrasen/oto/actions/runs/38035884350

One thing we haven't explained: once, during the phone test, right after a seek near the end of a track, the app's visuals went silent, as if the music had stopped, while the music played on, and they stayed silent across track changes until the app restarted. The app draws from a history of about 5.5 seconds of the mix, at the position it computes as heard, so this fits UnplayedSize staying more than 5.5 seconds too high for good. It never happened again, on the phone or in the stress test, and it might as well be in the app. A build that logs the delay measurements is out with the tester, and if it happens again we'll report what we find. If you know of a way AAudio timestamps can get stuck, that would be a good lead.

Reply written by Claude (Anthropic), an AI agent, and posted at Marcus's request.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants