Back to listening review
# Modal speech execution check — 25 September 2026

Two existing public-source clips ran once each on L4/CUDA using the pinned model and unchanged bounded launcher. No customer audio, deployment, integration, or audio playback.

| Language | Audio | Processing | Model loading | Predicted speakers |
|---|---:|---:|---:|---:|
| English | 10s | 1.63s | 24.59s | 2 |
| Arabic | 18s | 2.03s | 16.55s | 2 |

Processing excludes model loading, network transfer and platform startup. Input frames were replayed as fast as possible, so these are throughput measurements, not live listener delay. Both timelines complete with zero gaps/drops. All three successful GPU apps (including silence) stopped with zero tasks.

English source: https://www.youtube.com/watch?v=iixtaiov_nI at 330–340s. Arabic source: QatarDebate https://www.youtube.com/watch?v=Tp5yxE3JOOk at 1829–1847s. These correspond to listening-page clips 1 and 3.

Arabic GPU output marks speaker_1 at 5.05–5.81s and 13.05–13.27s. Human review must establish whether these are real interruptions or false speaker detections. No human reference, DER score, or accuracy claim. Short isolated excerpts also reset model context; direct comparison with the previous longer CPU runs is not an isolated GPU-versus-CPU accuracy comparison. The listening page shows both earlier CPU predictions and these GPU results for clips 1 and 3, clearly labelled.

Money: five conservative $1 reservations total ($5 of $20). This is not actual spend. Current workspace usage ceiling remains $1. Latest provider read at 04:26:36Z (after silence but before speech) showed $0.01 usage, $0 charges and $29.99 credits; billing can lag. Requested one post-speech refresh.

Observed billing refresh at 04:29:54Z: Cost Summary $0.02 total usage, covered by credits; $0.00 charged. Rounded component rows and the usage-limit widget differ because of rounding/reporting lag. This is the posted usage snapshot, not a final run cost. $5 reserved remains a separate conservative ledger figure.