Executive findings
A negative delta favors Wit because every delta is Wit minus Cohere. Confidence intervals that exclude zero are labelled decisive for this benchmark. Statistical significance does not by itself establish practical significance or generalization beyond these datasets.
Domain WER at a glance
July 12 production release
python transcribe_optimized.py --language ar --vad silero --vad-merge --alignment segment 1.wav. Static length-sorted batch 24 remains the RTX 3060 default. Upward adaptation and pinned host transfers remain opt-in because neither improved repeated end-to-end timing.| Configuration | External runs | Median wall | RTFx | Generation | Batches | Words | Transcript SHA-256 |
|---|---|---|---|---|---|---|---|
| Static batch 24, pageable (production default) | 36.58s / 36.39s | 36.485s | 114.04 | 22.370s | 8/8 | 9,852 | 11e07d260a45355f… |
| Adaptive batch 24→30 (experimental) | 36.82s / 36.40s | 36.610s | 113.65 | 22.358s | 7/7 | 9,852 | 51e635a60f44b98a… |
| Static batch 24, pinned host memory (experimental) | 36.29s / 36.61s | 36.450s | 114.15 | 22.129s | 8/8 | 9,852 | 11e07d260a45355f… |
The default median is based on 36.58s and 36.39s process-wall observations. The current pipeline used 182 merged model rows from 378 raw speech spans and retained 4080.224s of the 4160.679s recording. Its output contains 9,852 words and 1,198 subtitle cues. The adaptive transcript has the same word count but a different hash and 12 lexical token edits relative to static; batching is not assumed text-invariant under finite-precision GPU execution.
| Engine | Wall | RTFx | Original → compact cues | Words | Timing boundary |
|---|---|---|---|---|---|
| Cohere production default | 36.485s | 114.038 | 1,198 → 284 | 9,852 | Local RTX 3060; two-run wall midpoint |
| Wit.ai through stock Tafrigh 1.7.8 | 131.79s | 31.571 | 323 → 237 | 9,779 | Cloud; 8 independent Arabic apps; one observation |
This long lecture has no human transcript, so its 9,852 versus 9,779 output-word counts and audible disagreements are not WER. The comparison separates one-file latency from fleet throughput: 64 truly independent app quotas across eight Tafrigh processes can scale different files concurrently, while this 131.79-second measurement is one file through one eight-app process.
| Configuration | Wall | RTFx | Segments | Overall WER | MSA WER | Dialect WER | CA proxy WER | S/D/I |
|---|---|---|---|---|---|---|---|---|
| Prior rounded-boundary Silero (historical) | 48.70s | 103.34 | 729 | 32.4073% | 7.3939% | 51.4638% | 41.4723% | 1,039/696/353 |
| Sample-exact Silero, static 24 | 40.13s | 125.41 | 729 | 31.2898% | 7.5556% | 47.7273% | 43.0029% | 1,041/692/283 |
| Sample-exact Silero, adaptive | 40.15s | 125.35 | 729 | 31.0570% | 7.5556% | 47.6888% | 41.9825% | 1,036/695/270 |
| Sample-exact Silero + merge, static 24 | 41.50s | 121.27 | 508 | 27.6424% | 6.8687% | 46.9569% | 28.5714% | 852/738/191 |
| Stage | Measured time | Context |
|---|---|---|
| Decode worker | 4.523s | ffmpeg |
| Silero VAD worker | 5.918s | onnx on CPUExecutionProvider |
| ASR model load | 6.384s | Overlapped with audio preparation |
| ASR wall / generation | 23.223s / 22.368s | 8 batches; 28,277 generated tokens |
| Feature preparation wait | 0.828s | 5.703% aggregate frame padding |
| Host-to-device copies | 0.205s | Pageable host tensors in the production default |
| Area | Release behavior | Why it matters |
|---|---|---|
| Batch control | Persistent OOM cap learning and frame-balanced retry; upward growth stays opt-in | Prevents repeated oversized attempts without changing the measured default |
| Decoder completeness | EOS-aware token-limit detection with affected-row-only retry | Live two-token test retried one row at 128 and left no unresolved row |
| VAD | Sample-index timestamps, isolated ONNX/JIT loaders, provider/fallback provenance | Actual ONNX and JIT spans matched on a real clip |
| Approximate timing | Words distributed across retained raw speech spans | Merged silence is preserved as subtitle gaps instead of synthetic speech time |
| Alignment | FP32 accelerator-side stable log-softmax and in-place OOM reduction | Five-file real smoke test: 12 valid words, zero fallback segments |
| I/O | Concrete decode backend, streamed FFmpeg PCM, transactional profile/output publication | Avoids a full duplicate PCM buffer and ambiguous auto provenance |
| Compatibility | Transformers 5.13.x runtime bound and recorded validation environment | Internal model hooks now fail closed outside the validated series |
| Telemetry | Per-batch frames/tokens/padding/timing plus allocated and reserved VRAM | The remaining bottleneck can be measured rather than inferred from rounded logs |
The reviewer’s warning about TorchAudio forced_align removal was outdated: the operation was retained during the TorchAudio maintenance transition. Production therefore checks matching Torch/TorchAudio release lines and executes an operation smoke test rather than replacing a working dependency. The release passed 54 transcription-focused tests and 71 benchmark tests, Ruff lint/format, and bytecode compilation. Mypy was unavailable, so no mypy result is claimed.
The portable final/ bundle contains the production implementation as transcribe.py with simplified input.txt/input.srt/input.vtt/input.json output names, bundled vectorized Silero ONNX runtime and applicable licenses, portable runtime requirements, a device-oriented setup guide, version/model provenance, and an installation validator. Model weights remain upstream and download at first use after model-access approval.
Release summary artifact: production_release_20260712.json. Production audit: PRODUCTION_REVIEW_20260712.md. Both are hashed in the artifact inventory below.
Evaluation suite
The frozen suite contains one audio path and reference per clip. Exact waveform samples returned by the configured decoder, rather than source metadata alone, provide the timing denominator.
| Dataset | Role | Clips | Decoded h | Pinned revision | Source-card license |
|---|---|---|---|---|---|
| Casablanca | Eight conversational dialects | 6,726 | 7.703 | 8951b1b88e28 | CC BY-NC-ND 4.0 |
| Common Voice 18 Arabic | MSA / read speech | 10,471 | 12.657 | 1a52eefd8259 | CC0 |
| FLEURS ar-EG | MSA / read speech | 428 | 1.302 | 70bb2e84b976 | CC BY 4.0 |
| Quran ayah CA proxy | Classical Arabic recitation proxy | 600 | 3.982 | 80cad1ab411c | Treat as CC BY-NC-SA 4.0 |
| SADA22 | Saudi dialects; broadcast/noise/music | 6,189 | 10.749 | 094fe2c0fe4b | CC BY-NC-SA 4.0 |
Common Voice, SADA22, and Casablanca references were matched to the exact rows published by the Arabic ASR leaderboard at commit 10cf2c8. The Quran subset is 12 evenly spaced blocks of 50 rows from a three-reciter-held-out split. Its canonical verse references make it a Classical Arabic proxy, not an official Quranic leaderboard score.
Decoded duration was 36.393282 hours versus 36.301534 hours in source metadata, a 330.292-second difference dominated by Common Voice MP3 duration reporting. All 24,414 paths decoded successfully and IDs, references, grouping metadata, and audio paths matched across compared configurations.
Other Arabic evaluation options considered
No single corpus covers MSA, regional dialects, noise, code-switching, and Classical Arabic. These candidates remain useful extensions to the completed suite.
| Candidate | Test scale | Coverage | Status / tradeoff |
|---|---|---|---|
| Open Universal Arabic ASR Leaderboard | About 65.9 h | SADA, CV18, MASC, MGB-2, Casablanca | Best final public comparison; substantially larger and partly gated |
| MASC clean/noisy | About 25.4 h | 20+ dialects, MSA, YouTube, varied acoustics | Strong expansion; large download and not run here |
| MGB-2 | About 9.6 h | Al Jazeera broadcast, interviews, several dialects | Gated and research-only |
| ArzEn | About 2.9 h test | Spontaneous Egyptian Arabic-English code-switching | Access request required |
| TunSwitch-CS | About 25 min test | Tunisian Arabic-French-English code-switching | Open but small |
| Private in-domain holdout | Recommended 5-10 h | Actual production speakers, channels, and noise | Best production-validity check; split by speaker and source |
Full-suite accuracy
These are corpus-aggregated lexical-normalized rates. S/D/I means substitutions, deletions, and insertions. The bootstrap resamples 14,319 speaker/utterance clusters overall, with speakers namespaced by dataset where speaker IDs exist.
| Scope | Clips | Hours | Cohere WER | Wit WER | Delta | Paired 95% CI | Cohere CER | Wit CER | WER S/D/I Cohere / Wit |
Exact disagreements | Conclusion |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Overall | 24,414 | 36.39 | 31.32% | 34.00% | +2.68 pp | [+1.67, +4.16] pp | 14.24% | 20.03% | 46,576/17,756/7,629 34,499/38,952/4,669 |
23,178 / 94.94% | Cohere better |
| Scope | Clips | Hours | Cohere WER | Wit WER | Delta | Paired 95% CI | Cohere CER | Wit CER | WER S/D/I Cohere / Wit |
Exact disagreements | Conclusion |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Casablanca | 6,726 | 7.70 | 48.76% | 48.83% | +0.07 pp | [-0.48, +0.57] pp | 18.73% | 26.78% | 25,846/7,659/2,777 16,781/17,513/2,040 |
6,523 / 96.98% | No decisive difference |
| Common Voice 18 Arabic | 10,471 | 12.66 | 5.54% | 12.87% | +7.32 pp | [+6.71, +8.02] pp | 1.53% | 5.91% | 2,495/275/184 4,420/2,144/293 |
9,766 / 93.27% | Cohere better |
| FLEURS ar-EG | 428 | 1.30 | 4.75% | 19.73% | +14.99 pp | [+12.60, +17.52] pp | 2.15% | 14.68% | 232/96/52 534/909/137 |
428 / 100.00% | Cohere better |
| Quran ayah CA proxy | 600 | 3.98 | 15.35% | 41.70% | +26.35 pp | [+8.37, +53.13] pp | 11.07% | 38.46% | 252/50/914 329/2,835/140 |
593 / 98.83% | Cohere better |
| SADA22 | 6,189 | 10.75 | 36.14% | 34.89% | -1.26 pp | [-2.02, -0.50] pp | 20.00% | 21.90% | 17,751/9,676/3,702 12,435/15,551/2,059 |
5,868 / 94.81% | Wit better |
| Scope | Clips | Hours | Cohere WER | Wit WER | Delta | Paired 95% CI | Cohere CER | Wit CER | WER S/D/I Cohere / Wit |
Exact disagreements | Conclusion |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Classical Arabic | 600 | 3.98 | 15.35% | 41.70% | +26.35 pp | [+8.37, +53.13] pp | 11.07% | 38.46% | 252/50/914 329/2,835/140 |
593 / 98.83% | Cohere better |
| Dialect | 12,758 | 17.91 | 42.60% | 41.95% | -0.64 pp | [-1.12, -0.19] pp | 19.76% | 24.59% | 43,169/17,125/6,426 28,895/32,818/4,000 |
12,241 / 95.95% | Wit better |
| MSA | 11,056 | 14.50 | 6.17% | 13.96% | +7.79 pp | [+7.15, +8.47] pp | 1.95% | 7.23% | 3,155/581/289 5,275/3,299/529 |
10,344 / 93.56% | Cohere better |
Normalization-profile audit
Raw
Exact spelling, punctuation, whitespace-sensitive word boundaries, and diacritics. Useful for output fidelity, but sensitive to stylistic conventions.
Leaderboard repo-exact
Reproduces the public evaluator at commit 10cf2c8, including its malformed punctuation-regex behavior. Included as an audit profile.
Leaderboard intended
Applies the punctuation removal documented by the leaderboard plus Arabic diacritic, hamza, and digit mappings.
Lexical normalized
NFKC, combining-mark and tatweel removal, punctuation-to-space normalization, Arabic/Persian letter folding, and digit mapping. This is the headline profile and the closest disclosed approximation to Cohere's scoring.
| Scope | Profile | Cohere WER | Wit WER | Delta | Cohere CER | Wit CER | Ref words | Cohere WER S/D/I | Wit WER S/D/I |
|---|---|---|---|---|---|---|---|---|---|
| Overall | Raw | 49.09% | 48.85% | -0.24 pp | 21.76% | 26.24% | 230,422 | 86,285/18,039/8,797 | 68,504/39,408/4,650 |
| Overall | Leaderboard repo-exact | 38.03% | 38.04% | +0.01 pp | 16.59% | 21.13% | 230,421 | 60,615/18,128/8,887 | 43,428/39,487/4,730 |
| Overall | Leaderboard intended | 31.59% | 34.19% | +2.60 pp | 14.45% | 20.17% | 229,701 | 47,211/17,820/7,537 | 34,914/38,929/4,702 |
| Overall | Lexical normalized | 31.32% | 34.00% | +2.68 pp | 14.24% | 20.03% | 229,757 | 46,576/17,756/7,629 | 34,499/38,952/4,669 |
Every profile by dataset
| Scope | Profile | Cohere WER | Wit WER | Delta | Cohere CER | Wit CER | Ref words | Cohere WER S/D/I | Wit WER S/D/I |
|---|---|---|---|---|---|---|---|---|---|
| Casablanca | Raw | 58.09% | 59.33% | +1.25 pp | 22.58% | 29.78% | 75,085 | 32,796/8,141/2,677 | 24,462/18,033/2,056 |
| Casablanca | Leaderboard repo-exact | 54.55% | 55.50% | +0.95 pp | 21.24% | 28.64% | 75,085 | 30,068/8,178/2,714 | 21,527/18,061/2,084 |
| Casablanca | Leaderboard intended | 48.87% | 48.95% | +0.08 pp | 18.97% | 26.99% | 74,375 | 25,949/7,676/2,723 | 16,848/17,495/2,063 |
| Casablanca | Lexical normalized | 48.76% | 48.83% | +0.07 pp | 18.73% | 26.78% | 74,416 | 25,846/7,659/2,777 | 16,781/17,513/2,040 |
| Common Voice 18 Arabic | Raw | 38.12% | 43.13% | +5.02 pp | 15.73% | 18.59% | 53,289 | 18,976/276/1,060 | 20,550/2,145/291 |
| Common Voice 18 Arabic | Leaderboard repo-exact | 17.27% | 13.73% | -3.54 pp | 4.39% | 6.22% | 53,288 | 7,867/276/1,061 | 4,876/2,147/294 |
| Common Voice 18 Arabic | Leaderboard intended | 5.85% | 13.11% | +7.26 pp | 1.78% | 6.10% | 53,284 | 2,650/281/188 | 4,549/2,143/294 |
| Common Voice 18 Arabic | Lexical normalized | 5.54% | 12.87% | +7.32 pp | 1.53% | 5.91% | 53,286 | 2,495/275/184 | 4,420/2,144/293 |
| FLEURS ar-EG | Raw | 19.76% | 37.30% | +17.54 pp | 5.24% | 18.07% | 8,000 | 1,418/104/59 | 1,953/898/133 |
| FLEURS ar-EG | Leaderboard repo-exact | 15.34% | 21.84% | +6.50 pp | 4.17% | 15.11% | 8,000 | 1,064/104/59 | 690/911/146 |
| FLEURS ar-EG | Leaderboard intended | 4.90% | 19.81% | +14.91 pp | 2.22% | 14.74% | 7,994 | 239/98/55 | 533/905/146 |
| FLEURS ar-EG | Lexical normalized | 4.75% | 19.73% | +14.99 pp | 2.15% | 14.68% | 8,007 | 232/96/52 | 534/909/137 |
| Quran ayah CA proxy | Raw | 104.19% | 101.34% | -2.85 pp | 44.33% | 64.65% | 7,924 | 7,308/37/911 | 5,123/2,801/106 |
| Quran ayah CA proxy | Leaderboard repo-exact | 19.51% | 44.50% | +24.99 pp | 13.02% | 39.11% | 7,924 | 574/49/923 | 551/2,835/140 |
| Quran ayah CA proxy | Leaderboard intended | 19.07% | 44.18% | +25.11 pp | 11.92% | 39.05% | 7,924 | 550/50/911 | 526/2,835/140 |
| Quran ayah CA proxy | Lexical normalized | 15.35% | 41.70% | +26.35 pp | 11.07% | 38.46% | 7,924 | 252/50/914 | 329/2,835/140 |
| SADA22 | Raw | 45.70% | 39.49% | -6.21 pp | 23.33% | 23.02% | 86,124 | 25,787/9,481/4,090 | 16,416/15,531/2,064 |
| SADA22 | Leaderboard repo-exact | 40.28% | 38.76% | -1.52 pp | 21.92% | 22.84% | 86,124 | 21,042/9,521/4,130 | 15,784/15,533/2,066 |
| SADA22 | Leaderboard intended | 36.22% | 34.91% | -1.31 pp | 20.11% | 21.92% | 86,124 | 17,823/9,715/3,660 | 12,458/15,551/2,059 |
| SADA22 | Lexical normalized | 36.14% | 34.89% | -1.26 pp | 20.00% | 21.90% | 86,124 | 17,751/9,676/3,702 | 12,435/15,551/2,059 |
Every profile by domain
| Scope | Profile | Cohere WER | Wit WER | Delta | Cohere CER | Wit CER | Ref words | Cohere WER S/D/I | Wit WER S/D/I |
|---|---|---|---|---|---|---|---|---|---|
| Classical Arabic | Raw | 104.19% | 101.34% | -2.85 pp | 44.33% | 64.65% | 7,924 | 7,308/37/911 | 5,123/2,801/106 |
| Classical Arabic | Leaderboard repo-exact | 19.51% | 44.50% | +24.99 pp | 13.02% | 39.11% | 7,924 | 574/49/923 | 551/2,835/140 |
| Classical Arabic | Leaderboard intended | 19.07% | 44.18% | +25.11 pp | 11.92% | 39.05% | 7,924 | 550/50/911 | 526/2,835/140 |
| Classical Arabic | Lexical normalized | 15.35% | 41.70% | +26.35 pp | 11.07% | 38.46% | 7,924 | 252/50/914 | 329/2,835/140 |
| Dialect | Raw | 52.00% | 49.46% | -2.54 pp | 23.34% | 26.64% | 157,299 | 57,671/17,415/6,713 | 40,468/33,318/4,021 |
| Dialect | Leaderboard repo-exact | 47.56% | 47.25% | -0.31 pp | 21.97% | 26.00% | 157,299 | 50,532/17,491/6,789 | 36,923/33,348/4,051 |
| Dialect | Leaderboard intended | 42.69% | 42.03% | -0.67 pp | 19.93% | 24.70% | 156,589 | 43,341/17,181/6,330 | 28,985/32,800/4,023 |
| Dialect | Lexical normalized | 42.60% | 41.95% | -0.64 pp | 19.76% | 24.59% | 156,630 | 43,169/17,125/6,426 | 28,895/32,818/4,000 |
| MSA | Raw | 35.38% | 40.99% | +5.61 pp | 14.03% | 17.92% | 65,199 | 21,306/587/1,173 | 22,913/3,289/523 |
| MSA | Leaderboard repo-exact | 17.29% | 15.03% | -2.26 pp | 4.57% | 7.56% | 65,198 | 9,509/588/1,175 | 5,954/3,304/539 |
| MSA | Leaderboard intended | 6.45% | 14.17% | +7.72 pp | 2.16% | 7.39% | 65,188 | 3,320/589/296 | 5,403/3,294/539 |
| MSA | Lexical normalized | 6.17% | 13.96% | +7.79 pp | 1.95% | 7.23% | 65,203 | 3,155/581/289 | 5,275/3,299/529 |
Paired confidence intervals were computed only for the selected lexical-normalized profile. Deltas in the other profile tables are descriptive corpus-rate differences, not separately bootstrapped estimates.
Dialect and variety detail
Labels are retained from their source datasets. “MSA” is a reference-register label, “More than 1 speaker” and “Unknown” are annotation categories, and tiny categories should not be ranked.
| Scope | Clips | Hours | Cohere WER | Wit WER | Delta | Paired 95% CI | Cohere CER | Wit CER | WER S/D/I Cohere / Wit |
Exact disagreements | Conclusion |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Algeria | 843 | 0.94 | 63.04% | 59.53% | -3.51 pp | [-5.49, -1.82] pp | 24.31% | 31.71% | 3,702/530/760 2,487/1,759/468 |
825 / 97.86% | Wit better |
| Classical Arabic | 600 | 3.98 | 15.35% | 41.70% | +26.35 pp | [+8.37, +53.13] pp | 11.07% | 38.46% | 252/50/914 329/2,835/140 |
593 / 98.83% | Cohere better |
| Egypt | 825 | 0.99 | 31.61% | 36.69% | +5.07 pp | [+3.64, +6.41] pp | 11.93% | 15.46% | 2,120/863/208 2,054/1,283/366 |
781 / 94.67% | Cohere better |
| Egyptian | 96 | 0.09 | 35.24% | 30.64% | -4.61 pp | [-10.26, +0.52] pp | 16.74% | 14.54% | 185/71/27 151/65/30 |
85 / 88.54% | No decisive difference |
| Hijazi | 809 | 1.12 | 31.99% | 32.90% | +0.91 pp | [-0.71, +2.60] pp | 15.96% | 20.78% | 1,816/802/249 1,180/1,613/156 |
773 / 95.55% | No decisive difference |
| Jordan | 848 | 0.98 | 29.69% | 31.93% | +2.24 pp | [+1.01, +3.42] pp | 8.82% | 11.96% | 1,894/569/147 1,705/912/190 |
809 / 95.40% | Cohere better |
| Khaliji | 1,150 | 1.13 | 37.70% | 42.20% | +4.49 pp | [+2.17, +6.64] pp | 18.56% | 28.57% | 2,257/829/445 1,435/2,338/179 |
1,081 / 94.00% | Cohere better |
| MSA | 11,056 | 14.50 | 6.17% | 13.96% | +7.79 pp | [+7.14, +8.48] pp | 1.95% | 7.23% | 3,155/581/289 5,275/3,299/529 |
10,344 / 93.56% | Cohere better |
| Mauritania | 948 | 0.94 | 79.75% | 84.94% | +5.19 pp | [+3.81, +6.56] pp | 41.53% | 73.05% | 5,481/2,060/445 1,408/7,016/82 |
940 / 99.16% | Cohere better |
| More than 1 speaker اكثر من متحدث | 1,320 | 4.80 | 39.42% | 35.86% | -3.56 pp | [-4.92, -2.32] pp | 23.46% | 22.22% | 7,769/5,682/1,794 6,227/6,527/1,115 |
1,303 / 98.71% | Wit better |
| Morocco | 1,045 | 1.01 | 54.50% | 43.01% | -11.50 pp | [-12.78, -10.23] pp | 17.41% | 17.77% | 4,856/1,382/280 2,900/2,084/159 |
1,020 / 97.61% | Wit better |
| Najdi | 1,703 | 2.07 | 31.04% | 28.31% | -2.73 pp | [-4.11, -1.45] pp | 16.49% | 16.54% | 3,288/1,499/553 2,096/2,436/338 |
1,571 / 92.25% | Wit better |
| Notapplicable | 167 | 0.13 | 44.12% | 51.02% | +6.90 pp | [+2.91, +10.93] pp | 22.36% | 34.82% | 264/86/59 172/292/9 |
158 / 94.61% | Cohere better |
| Palestine | 667 | 0.98 | 37.96% | 36.26% | -1.69 pp | [-2.86, -0.54] pp | 12.41% | 12.87% | 2,358/746/241 2,010/918/268 |
645 / 96.70% | Wit better |
| Shamali | 18 | 0.02 | 26.73% | 30.69% | +3.96 pp | [-0.91, +12.50] pp | 10.62% | 18.28% | 47/4/3 35/25/2 |
16 / 88.89% | No decisive difference |
| UAE | 813 | 0.93 | 40.38% | 45.19% | +4.81 pp | [+3.29, +6.28] pp | 12.92% | 24.30% | 2,429/799/228 1,829/1,863/176 |
789 / 97.05% | Cohere better |
| Unknown | 762 | 0.83 | 44.40% | 48.54% | +4.14 pp | [+0.28, +7.58] pp | 25.57% | 36.35% | 1,680/482/517 800/2,003/126 |
724 / 95.01% | Cohere better |
| Yemen | 737 | 0.94 | 50.62% | 53.19% | +2.58 pp | [+1.27, +3.88] pp | 19.62% | 27.42% | 3,006/710/468 2,388/1,678/331 |
714 / 96.88% | Cohere better |
| Yemeni | 7 | 0.01 | 65.22% | 63.04% | -2.17 pp | [-18.75, +9.43] pp | 59.09% | 27.73% | 17/11/2 18/6/5 |
7 / 100.00% | No decisive difference |
Every scoring profile for every dialect label
| Scope | Profile | Cohere WER | Wit WER | Delta | Cohere CER | Wit CER | Ref words | Cohere WER S/D/I | Wit WER S/D/I |
|---|---|---|---|---|---|---|---|---|---|
| Algeria | Raw | 66.76% | 65.20% | -1.55 pp | 28.29% | 33.64% | 8,044 | 4,023/633/714 | 2,957/1,838/450 |
| Algeria | Leaderboard repo-exact | 65.10% | 63.80% | -1.31 pp | 27.46% | 33.21% | 8,044 | 3,888/634/715 | 2,836/1,842/454 |
| Algeria | Leaderboard intended | 62.66% | 59.51% | -3.15 pp | 24.47% | 31.96% | 7,914 | 3,694/534/731 | 2,488/1,754/468 |
| Algeria | Lexical normalized | 63.04% | 59.53% | -3.51 pp | 24.31% | 31.71% | 7,919 | 3,702/530/760 | 2,487/1,759/468 |
| Classical Arabic | Raw | 104.19% | 101.34% | -2.85 pp | 44.33% | 64.65% | 7,924 | 7,308/37/911 | 5,123/2,801/106 |
| Classical Arabic | Leaderboard repo-exact | 19.51% | 44.50% | +24.99 pp | 13.02% | 39.11% | 7,924 | 574/49/923 | 551/2,835/140 |
| Classical Arabic | Leaderboard intended | 19.07% | 44.18% | +25.11 pp | 11.92% | 39.05% | 7,924 | 550/50/911 | 526/2,835/140 |
| Classical Arabic | Lexical normalized | 15.35% | 41.70% | +26.35 pp | 11.07% | 38.46% | 7,924 | 252/50/914 | 329/2,835/140 |
| Egypt | Raw | 44.93% | 51.00% | +6.07 pp | 16.17% | 19.41% | 10,116 | 3,448/874/223 | 3,494/1,292/373 |
| Egypt | Leaderboard repo-exact | 39.83% | 45.35% | +5.53 pp | 14.73% | 17.86% | 10,116 | 2,928/876/225 | 2,907/1,300/381 |
| Egypt | Leaderboard intended | 32.50% | 37.08% | +4.58 pp | 12.33% | 15.65% | 10,093 | 2,211/862/207 | 2,094/1,282/366 |
| Egypt | Lexical normalized | 31.61% | 36.69% | +5.07 pp | 11.93% | 15.46% | 10,094 | 2,120/863/208 | 2,054/1,283/366 |
| Egyptian | Raw | 42.09% | 35.37% | -6.72 pp | 18.82% | 15.79% | 803 | 242/70/26 | 187/65/32 |
| Egyptian | Leaderboard repo-exact | 38.85% | 34.99% | -3.86 pp | 17.82% | 15.68% | 803 | 216/70/26 | 184/65/32 |
| Egyptian | Leaderboard intended | 36.49% | 30.76% | -5.73 pp | 17.10% | 14.59% | 803 | 195/71/27 | 152/65/30 |
| Egyptian | Lexical normalized | 35.24% | 30.64% | -4.61 pp | 16.74% | 14.54% | 803 | 185/71/27 | 151/65/30 |
| Hijazi | Raw | 41.07% | 36.64% | -4.43 pp | 19.02% | 21.66% | 8,963 | 2,625/795/261 | 1,519/1,611/154 |
| Hijazi | Leaderboard repo-exact | 34.90% | 35.83% | +0.93 pp | 17.34% | 21.48% | 8,963 | 2,062/800/266 | 1,444/1,612/155 |
| Hijazi | Leaderboard intended | 32.11% | 32.91% | +0.80 pp | 16.03% | 20.79% | 8,963 | 1,827/805/246 | 1,181/1,613/156 |
| Hijazi | Lexical normalized | 31.99% | 32.90% | +0.91 pp | 15.96% | 20.78% | 8,963 | 1,816/802/249 | 1,180/1,613/156 |
| Jordan | Raw | 40.16% | 45.77% | +5.61 pp | 11.87% | 15.66% | 8,804 | 2,801/577/158 | 2,919/912/199 |
| Jordan | Leaderboard repo-exact | 34.23% | 37.89% | +3.66 pp | 10.09% | 13.51% | 8,804 | 2,273/580/161 | 2,217/916/203 |
| Jordan | Leaderboard intended | 29.75% | 31.94% | +2.18 pp | 8.90% | 12.02% | 8,792 | 1,899/570/147 | 1,706/912/190 |
| Jordan | Lexical normalized | 29.69% | 31.93% | +2.24 pp | 8.82% | 11.96% | 8,792 | 1,894/569/147 | 1,705/912/190 |
| Khaliji | Raw | 48.64% | 44.49% | -4.15 pp | 23.03% | 29.17% | 9,366 | 3,245/808/503 | 1,653/2,335/179 |
| Khaliji | Leaderboard repo-exact | 41.22% | 43.92% | +2.70 pp | 20.97% | 29.02% | 9,366 | 2,540/813/508 | 1,598/2,336/180 |
| Khaliji | Leaderboard intended | 37.84% | 42.28% | +4.44 pp | 18.72% | 28.60% | 9,366 | 2,271/830/443 | 1,443/2,338/179 |
| Khaliji | Lexical normalized | 37.70% | 42.20% | +4.49 pp | 18.56% | 28.57% | 9,366 | 2,257/829/445 | 1,435/2,338/179 |
| MSA | Raw | 35.38% | 40.99% | +5.61 pp | 14.03% | 17.92% | 65,199 | 21,306/587/1,173 | 22,913/3,289/523 |
| MSA | Leaderboard repo-exact | 17.29% | 15.03% | -2.26 pp | 4.57% | 7.56% | 65,198 | 9,509/588/1,175 | 5,954/3,304/539 |
| MSA | Leaderboard intended | 6.45% | 14.17% | +7.72 pp | 2.16% | 7.39% | 65,188 | 3,320/589/296 | 5,403/3,294/539 |
| MSA | Lexical normalized | 6.17% | 13.96% | +7.79 pp | 1.95% | 7.23% | 65,203 | 3,155/581/289 | 5,275/3,299/529 |
| Mauritania | Raw | 82.29% | 86.61% | +4.31 pp | 44.85% | 73.70% | 10,110 | 5,774/2,130/416 | 1,566/7,108/82 |
| Mauritania | Leaderboard repo-exact | 81.22% | 85.87% | +4.65 pp | 43.59% | 73.43% | 10,110 | 5,651/2,137/423 | 1,491/7,108/82 |
| Mauritania | Leaderboard intended | 79.66% | 84.97% | +5.31 pp | 41.67% | 73.11% | 10,004 | 5,475/2,071/423 | 1,410/7,007/83 |
| Mauritania | Lexical normalized | 79.75% | 84.94% | +5.19 pp | 41.53% | 73.05% | 10,014 | 5,481/2,060/445 | 1,408/7,016/82 |
| More than 1 speaker اكثر من متحدث | Raw | 48.53% | 42.46% | -6.07 pp | 26.80% | 23.83% | 38,672 | 11,204/5,539/2,025 | 8,782/6,518/1,121 |
| More than 1 speaker اكثر من متحدث | Leaderboard repo-exact | 44.79% | 41.74% | -3.05 pp | 25.81% | 23.66% | 38,672 | 9,721/5,557/2,043 | 8,504/6,518/1,121 |
| More than 1 speaker اكثر من متحدث | Leaderboard intended | 39.50% | 35.88% | -3.62 pp | 23.58% | 22.24% | 38,672 | 7,786/5,698/1,790 | 6,234/6,527/1,115 |
| More than 1 speaker اكثر من متحدث | Lexical normalized | 39.42% | 35.86% | -3.56 pp | 23.46% | 22.22% | 38,672 | 7,769/5,682/1,794 | 6,227/6,527/1,115 |
| Morocco | Raw | 63.23% | 53.95% | -9.28 pp | 20.78% | 20.73% | 11,959 | 5,921/1,374/267 | 4,184/2,076/192 |
| Morocco | Leaderboard repo-exact | 62.20% | 52.17% | -10.03 pp | 20.08% | 20.05% | 11,959 | 5,794/1,376/269 | 3,971/2,076/192 |
| Morocco | Leaderboard intended | 54.43% | 43.01% | -11.41 pp | 17.42% | 17.83% | 11,959 | 4,855/1,386/268 | 2,901/2,084/159 |
| Morocco | Lexical normalized | 54.50% | 43.01% | -11.50 pp | 17.41% | 17.77% | 11,959 | 4,856/1,382/280 | 2,900/2,084/159 |
| Najdi | Raw | 40.64% | 31.57% | -9.07 pp | 19.37% | 17.29% | 17,201 | 4,921/1,483/587 | 2,662/2,432/337 |
| Najdi | Leaderboard repo-exact | 33.99% | 30.59% | -3.40 pp | 17.71% | 17.05% | 17,201 | 3,758/1,492/596 | 2,493/2,432/337 |
| Najdi | Leaderboard intended | 31.02% | 28.34% | -2.69 pp | 16.55% | 16.55% | 17,201 | 3,292/1,512/532 | 2,100/2,436/338 |
| Najdi | Lexical normalized | 31.04% | 28.31% | -2.73 pp | 16.49% | 16.54% | 17,201 | 3,288/1,499/553 | 2,096/2,436/338 |
| Notapplicable | Raw | 49.84% | 53.29% | +3.45 pp | 24.05% | 35.46% | 927 | 316/84/62 | 193/292/9 |
| Notapplicable | Leaderboard repo-exact | 45.85% | 52.97% | +7.12 pp | 22.90% | 35.40% | 927 | 279/84/62 | 190/292/9 |
| Notapplicable | Leaderboard intended | 44.55% | 51.13% | +6.58 pp | 22.47% | 34.86% | 927 | 268/86/59 | 173/292/9 |
| Notapplicable | Lexical normalized | 44.12% | 51.02% | +6.90 pp | 22.36% | 34.82% | 927 | 264/86/59 | 172/292/9 |
| Palestine | Raw | 47.99% | 51.12% | +3.12 pp | 15.36% | 16.95% | 8,833 | 3,254/752/233 | 3,314/920/281 |
| Palestine | Leaderboard repo-exact | 43.51% | 44.62% | +1.11 pp | 13.97% | 15.06% | 8,833 | 2,844/759/240 | 2,732/924/285 |
| Palestine | Leaderboard intended | 37.96% | 36.30% | -1.66 pp | 12.50% | 13.00% | 8,812 | 2,358/746/241 | 2,014/917/268 |
| Palestine | Lexical normalized | 37.96% | 36.26% | -1.69 pp | 12.41% | 12.87% | 8,813 | 2,358/746/241 | 2,010/918/268 |
| Shamali | Raw | 36.14% | 33.66% | -2.48 pp | 13.21% | 18.80% | 202 | 63/4/6 | 41/25/2 |
| Shamali | Leaderboard repo-exact | 30.69% | 31.19% | +0.50 pp | 11.85% | 18.49% | 202 | 52/4/6 | 36/25/2 |
| Shamali | Leaderboard intended | 26.73% | 30.69% | +3.96 pp | 10.73% | 18.28% | 202 | 47/4/3 | 35/25/2 |
| Shamali | Lexical normalized | 26.73% | 30.69% | +3.96 pp | 10.62% | 18.28% | 202 | 47/4/3 | 35/25/2 |
| UAE | Raw | 52.70% | 59.02% | +6.32 pp | 17.52% | 28.54% | 8,766 | 3,428/969/223 | 2,969/2,045/160 |
| UAE | Leaderboard repo-exact | 49.41% | 53.35% | +3.95 pp | 16.29% | 26.84% | 8,766 | 3,131/973/227 | 2,460/2,051/166 |
| UAE | Leaderboard intended | 40.38% | 45.24% | +4.86 pp | 13.20% | 24.54% | 8,559 | 2,435/799/222 | 1,833/1,863/176 |
| UAE | Lexical normalized | 40.38% | 45.19% | +4.81 pp | 12.92% | 24.30% | 8,559 | 2,429/799/228 | 1,829/1,863/176 |
| Unknown | Raw | 54.41% | 50.98% | -3.43 pp | 29.76% | 37.01% | 6,034 | 2,238/481/564 | 949/2,001/126 |
| Unknown | Leaderboard repo-exact | 47.51% | 50.61% | +3.10 pp | 27.81% | 36.92% | 6,034 | 1,818/483/566 | 927/2,001/126 |
| Unknown | Leaderboard intended | 44.45% | 48.56% | +4.11 pp | 25.78% | 36.38% | 6,034 | 1,689/488/505 | 801/2,003/126 |
| Unknown | Lexical normalized | 44.40% | 48.54% | +4.14 pp | 25.57% | 36.35% | 6,034 | 1,680/482/517 | 800/2,003/126 |
| Yemen | Raw | 64.14% | 61.75% | -2.39 pp | 25.53% | 30.67% | 8,453 | 4,147/832/443 | 3,059/1,842/319 |
| Yemen | Leaderboard repo-exact | 57.45% | 60.07% | +2.63 pp | 23.13% | 30.07% | 8,453 | 3,559/843/454 | 2,913/1,844/321 |
| Yemen | Leaderboard intended | 51.13% | 53.76% | +2.63 pp | 20.43% | 28.19% | 8,242 | 3,022/708/484 | 2,402/1,676/353 |
| Yemen | Lexical normalized | 50.62% | 53.19% | +2.58 pp | 19.62% | 27.42% | 8,266 | 3,006/710/468 | 2,388/1,678/331 |
| Yemeni | Raw | 71.74% | 67.39% | -4.35 pp | 62.73% | 28.64% | 46 | 21/10/2 | 20/6/5 |
| Yemeni | Leaderboard repo-exact | 65.22% | 67.39% | +2.17 pp | 60.00% | 28.64% | 46 | 18/10/2 | 20/6/5 |
| Yemeni | Leaderboard intended | 65.22% | 63.04% | -2.17 pp | 59.55% | 27.73% | 46 | 17/11/2 | 18/6/5 |
| Yemeni | Lexical normalized | 65.22% | 63.04% | -2.17 pp | 59.09% | 27.73% | 46 | 17/11/2 | 18/6/5 |
Timing and throughput
| Engine/config | Wall | Decoded RTFx | Clips/s | Scope |
|---|---|---|---|---|
| Optimized Cohere BF16 length batch 24 | 28m 25s | 76.82 | 14.31 | Local GPU, all-samples wall |
| Wit.ai via Tafrigh default | 1h 13m 17s | 29.79 | 5.55 | Cloud API + Tafrigh preprocessing, cumulative active wall |
| Component | Time | Interpretation |
|---|---|---|
| Cohere audio decode | 16m 08s | Serialized aggregate component |
| Cohere feature processing | 0m 50s | Serialized aggregate component |
| Cohere model generation | 11m 21s | Generation-only RTFx 192.25; not equivalent to Wit wall |
| Cohere text decode | 0.425s | Aggregate |
| Wit API request time sum | 20h 42m 33s | Sum across concurrent requests; exceeds wall by design |
| Wit preprocessing worker time sum | 1h 53m 09s | Sum across concurrent workers; not wall |
| Wit active wall | 1h 13m 17s | Recovered invocations combined |
Cohere generation alone was 11m 21s, but excluding decode and feature preparation makes it an invalid direct boundary against Wit end-to-end wall. Wit active time excludes credential entry, validation, debugging downtime, exact-duration analysis, and WER scoring. Its first recovered component is accurate to roughly one second; later components came from sanitized ledgers.
Continuous 1.wav case study
| Run | Evidence | Wall | RTFx | Key stages | Output |
|---|---|---|---|---|---|
| July 12 production: Silero + merge + segment | Two /usr/bin/time runs + exact profiles | 36.485s median | 114.04 | 22.37s generation; 8 batches; 6.06/6.10 GiB allocated/reserved | 9,852 words; 1,198 cues |
| Original transcribe.py | Exact stdout | 2m 38s | 26.3 | 62s ASR; 44s emissions/alignment | 9,754 words; 1,225 cues |
| July 11 historical: plain transcript only | Exact stdout + /usr/bin/time | 50.61s external | 82.2 | 24s ASR; alignment/cues skipped; 6.06 GiB allocated peak | 9,755 words; 378 text lines |
| July 11 historical: segment interpolation | Exact stdout + /usr/bin/time | 49.23s external | 84.5 | 23s ASR; alignment skipped; 6.06 GiB allocated peak | 9,755 words; 1,131 cues |
| July 11 historical: FP16 word alignment | Exact stdout + /usr/bin/time | 68.41s external | 60.8 | 23s ASR; 16s emissions; 6.06 GiB allocated peak | 9,755 words; 1,225 cues |
| July 11 historical: FP32 word alignment | Exact stdout + /usr/bin/time | 97.56s external | 42.7 | 23s ASR; 46s emissions; 6.06 GiB allocated peak | 9,755 words; 1,225 cues |
| Native F16, E4/D8 argmax | Guarded release harness | 84.36s external | 49.3 | 36.9s encoder; 40.6s decoder; 6,900 MiB peak | 9,881 words; text only |
| Native Q8_0, E4/D8 argmax | Guarded release harness | 60.55s external | 68.7 | 31.3s encoder; 22.6s decoder; 5,144 MiB peak | 9,881 words; text only |
| Native Q4_K imatrix, E4/D8 argmax | Guarded release harness | 53.17s external | 78.3 | 31.8s encoder; 14.9s decoder; 4,264 MiB peak | 9,882 words; text only |
In the frozen July 11 timestamp experiment, all four output/timestamp modes shared the same ASR inputs and produced the same joined text. That statement does not compare later sample-exact or merged segmentation policies. --text-only writes one non-empty VAD/ASR segment per line and skips word allocation, cue grouping, and MMS alignment; its historical 50.61-second wall was effectively tied with segment mode because ASR and VAD dominated. A frozen, word-indexed rerun paired 9,755 words between FP16 and FP32. The combined-boundary absolute median/p95/p99 were 0.00/0.00/0.00 ms, and the maximum was 1420.00 ms. This supersedes the earlier ordinal-SRT comparison: cue reflow changed which words an ordinal cue contained, so its apparent 3.12-second maximum was not a valid acoustic-boundary comparison. FP32 remains the implementation-reference word-alignment mode; --alignment segment is the production reliable/fast recommendation when exact word boundaries are unnecessary.
The native rows use the reference quiet-boundary planner rather than Silero and write no timestamps. Native Q8 differed from native F16 by 28 lexical word edits (0.2834%); Q4 differed by 116 (1.1740%). Because 1.wav has no human transcript, these are model disagreements, not WER. The frozen 500-clip reference benchmark, not this file, determines the precision recommendation.
VAD and text-only segmentation modes
--text-only removes timestamp and subtitle work; --vad controls which waveform regions reach the ASR model. Changing VAD can therefore change recognition text and accuracy, even though every mode writes the same plain-text output format. The production default remains 30-second maximum spans. The pinned processor accepts one row through 35 seconds, and the script verifies the actual processor row count at runtime; 35 seconds is supported as an experiment, not presumed faster or more accurate.
| Mode | Options | External wall | RTFx | Segments | Audio retained | Words | Normalized distance vs Silero |
|---|---|---|---|---|---|---|---|
| Silero ONNX | --text-only --vad silero |
50.61s | 82.21 | 378 | 95.75% | 9,755 | Reference mode |
| Auditok, threshold 50 | --text-only --vad auditok --energy-threshold 50 |
36.45s | 114.15 | 144 | 99.95% | 9,863 | 9.0909% (886 edits) |
| No VAD, fixed 30 s | --text-only --vad none --max-dur 30 |
34.97s | 118.98 | 139 | 100.00% | 9,905 | 9.1627% (893 edits) |
1.wav has no human reference transcript. Auditok and no-VAD were 9.09% and 9.16% different from Silero by normalized token edit distance. This roughly 9% result is disagreement, not accuracy, and does not identify which transcript is correct.| Mode | External wall | RTFx | Run VAD worker | Standalone VAD audit | Segments | Audio retained | No segments | Lexical WER | WER delta vs Silero | S/D/I |
|---|---|---|---|---|---|---|---|---|---|---|
| Silero ONNX | 48.70s | 103.40 | 27s | 21.49s | 729 | 69.25% | 14 | 32.4073% | Baseline | 1,039/696/353 |
| Auditok, threshold 50 | 49.75s | 101.22 | 4s | 1.71s | 723 | 88.79% | 1 | 28.4184% | -3.9888 pp | 948/537/346 |
| No VAD, fixed 30 s | 49.58s | 101.57 | 0s | 0.00s | 542 | 100.00% | 0 | 23.2656% | -9.1417 pp | 835/384/280 |
The same already-segmented corpus scored 22.9551% WER on the direct ASR benchmark path, outside this end-to-end script timing boundary. Silero completed in 48.70s, Auditok in 49.75s, and no-VAD in 49.58s. The cheaper VAD stages did not reduce complete folder wall in this single observation because preprocessing overlaps model loading/GPU work and each segmentation policy changed the ASR inputs.
| Dataset/domain | Silero WER | Auditok 50 WER | No-VAD 30 s WER | Auditok delta | No-VAD delta |
|---|---|---|---|---|---|
| Casablanca | 55.8099% | 52.2007% | 50.7923% | -3.6092 pp | -5.0176 pp |
| Common Voice 18 Arabic | 5.4688% | 8.2031% | 5.4688% | +2.7344 pp | +0.0000 pp |
| FLEURS ar-EG | 7.8961% | 17.2695% | 6.6225% | +9.3734 pp | -1.2736 pp |
| Quran ayah CA proxy | 41.4723% | 19.0962% | 15.1603% | -22.3761 pp | -26.3120 pp |
| SADA22 | 48.0822% | 40.7534% | 38.0822% | -7.3288 pp | -10.0000 pp |
| Dataset/domain | Silero WER | Auditok 50 WER | No-VAD 30 s WER | Auditok delta | No-VAD delta |
|---|---|---|---|---|---|
| Classical Arabic | 41.4723% | 19.0962% | 15.1603% | -22.3761 pp | -26.3120 pp |
| Dialect | 51.4638% | 45.7627% | 43.6441% | -5.7011 pp | -7.8197 pp |
| MSA | 7.3939% | 15.3939% | 6.3838% | +8.0000 pp | -1.0101 pp |
Silero remains the accuracy-conservative default for arbitrary continuous audio. Auditok threshold 50 and --vad none --max-dur 30 are explicit speed experiments whose suitability should be measured on representative, continuously recorded ground truth.
Timestamp-mode benchmark
| Mode | External wall | RTFx | Wall vs segment | CTC emissions | Words | Cues | Fallback segments | Lexical WER |
|---|---|---|---|---|---|---|---|---|
| Plain transcript only | 48.70s | 103.40 | 1.00× | Skipped | 6,119 | N/A | 0 | 32.4073% |
| Segment interpolation | 48.52s | 103.79 | 1.00× | Skipped | 6,119 | 1,032 | 0 | 32.4073% |
| FP16 forced word alignment | 119.07s | 42.29 | 2.45× | 66.00s | 6,119 | 1,194 | 2 | 32.4073% |
| FP32 forced word alignment | 233.26s | 21.59 | 4.81× | 177.00s | 6,119 | 1,192 | 2 | 32.4073% |
--text-only completed in 48.70s versus 48.52s for segment timing: a 0.18-second difference inside run-to-run variation. It produces only .aligned.txt, with one non-empty VAD/ASR segment per line. The periodic collection fix discovered during this benchmark reduced FP16 wall from 192.24s to 119.07s (38.1%) and FP32 from 302.15s to 233.26s (22.8%). Full heap collection now runs every 64 aligned files instead of after every file.
| Mode | Shared decode | Model load | Timing compute | Cold equivalent | Compute RTFx | Alignment peak GiB |
|---|---|---|---|---|---|---|
| Segment interpolation | 20.23s | 0.00s | 0.01s | 20.26s | 399167.81 | N/A |
| FP16 forced word alignment | 20.23s | 2.03s | 66.64s | 88.92s | 75.56 | 2.26 |
| FP32 forced word alignment | 20.23s | 1.60s | 178.00s | 199.85s | 28.29 | 4.10 |
On identical frozen audio, boundaries, and ASR text, FP16 timing compute was 2.67× faster than FP32. The complete production ratio is smaller because both modes pay the same ASR, VAD, startup, decode, and output costs.
| Comparison | Abs mean ms | Median | p95 | p99 | Max | ≤20 ms | ≤100 ms | >500 ms | >1 s | Exact word intervals | Matched cue spans |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Segment interpolation vs FP32 | 414.05 | 260.00 | 1331.85 | 2289.20 | 7760.00 | 12.706% | 27.341% | 28.1255% | 9.5604% | 1.59% | 610/1,192 |
| FP16 CTC vs FP32 | 0.91 | 0.00 | 0.00 | 0.00 | 1740.00 | 99.665% | 99.853% | 0.0327% | 0.0163% | 99.10% | 1,190/1,192 |
FP16 changed only a small tail: 99.665% of word boundaries were within 20 ms and 99.853% were within 100 ms, but four boundaries exceeded 500 ms and two exceeded one second. Both one-second outliers were in the Quran proxy, where a repeated phrase selected a different CTC path. Median, p95, and p99 alone would hide that case. Segment timing is intentionally approximate: it spreads words uniformly inside each VAD segment.
| Dataset | Segment median | Segment p95 | Segment max | FP16 median | FP16 p95 | FP16 p99 | FP16 max | FP16 >500 ms |
|---|---|---|---|---|---|---|---|---|
| Common Voice 18 Arabic | 180.00 | 560.00 | 1410.00 | 0.00 | 0.00 | 0.00 | 60.00 | 0 |
| FLEURS ar-EG | 266.67 | 970.96 | 7760.00 | 0.00 | 0.00 | 0.00 | 60.00 | 0 |
| SADA22 | 240.00 | 1403.86 | 6370.00 | 0.00 | 0.00 | 0.00 | 80.00 | 0 |
| Casablanca | 172.62 | 796.81 | 2640.00 | 0.00 | 0.00 | 0.00 | 80.00 | 0 |
| Quran ayah CA proxy | 580.00 | 1860.01 | 3546.67 | 0.00 | 0.00 | 60.00 | 1740.00 | 4 |
| Domain | Segment median | Segment p95 | Segment max | FP16 median | FP16 p95 | FP16 p99 | FP16 max | FP16 >500 ms |
|---|---|---|---|---|---|---|---|---|
| MSA | 244.88 | 920.00 | 7760.00 | 0.00 | 0.00 | 0.00 | 60.00 | 0 |
| Dialect | 206.67 | 1126.80 | 6370.00 | 0.00 | 0.00 | 0.00 | 80.00 | 0 |
| Classical Arabic | 580.00 | 1860.01 | 3546.67 | 0.00 | 0.00 | 60.00 | 1740.00 | 4 |
| Dataset/domain | Text-only WER | Segment WER | FP16 WER | FP32 WER |
|---|---|---|---|---|
| Common Voice 18 Arabic | 5.4688% | 5.4688% | 5.4688% | 5.4688% |
| FLEURS ar-EG | 7.8961% | 7.8961% | 7.8961% | 7.8961% |
| SADA22 | 48.0822% | 48.0822% | 48.0822% | 48.0822% |
| Casablanca | 55.8099% | 55.8099% | 55.8099% | 55.8099% |
| Quran ayah CA proxy | 41.4723% | 41.4723% | 41.4723% | 41.4723% |
| MSA | 7.3939% | 7.3939% | 7.3939% | 7.3939% |
| Dialect | 51.4638% | 51.4638% | 51.4638% | 51.4638% |
| Classical Arabic | 41.4723% | 41.4723% | 41.4723% | 41.4723% |
All 500 hypotheses and normalized texts matched across all four modes; the three timing modes also matched every structured word key. FP16 and FP32 each used uniform fallback timing for the same 97 words in two pathological repetition segments; fallback rows are included in the overall comparison and separately identified in the machine-readable report. Cue timing was matched by underlying word span, never by cue ordinal.
Cohere optimization evidence
Four complete BF16 batch/order configurations were evaluated on the same 24,414 clips. Length-sorted batch 24 was selected as the foundation. The later projection cache and conservative repetition-loop stop produced the final Cohere result used throughout the engine comparison.
| Config | Order | Batch | Wall | Wall RTFx | Generation | Gen RTFx | WER | CER | Δ vs ordered | Paired 95% CI | Peak GiB | OOM splits |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bf16_ordered_b24 | ordered | 24 | 48m 54s | 44.66 | 28m 44s | 76.01 | 32.301% | 16.429% | Baseline | Not estimated | 9.75 | 0 |
| bf16_length_b24 | length | 24 | 33m 26s | 65.32 | 14m 51s | 147.00 | 32.258% | 16.386% | -0.043 pp | [-0.21, +0.08] pp | 10.18 | 1 |
| bf16_length_b32 | length | 32 | 33m 48s | 64.60 | 15m 12s | 143.69 | 32.413% | 16.546% | +0.112 pp | [-0.14, +0.35] pp | 11.03 | 1 |
| bf16_length_b16 | length | 16 | 34m 11s | 63.87 | 15m 42s | 139.07 | 32.284% | 16.423% | -0.017 pp | [-0.29, +0.22] pp | 9.85 | 0 |
| Config | Wall | Wall RTFx | Generation | Gen RTFx | WER | CER | WER delta | Paired 95% CI |
|---|---|---|---|---|---|---|---|---|
| bf16_length_b24 | 33m 26s | 65.32 | 14m 51s | 147.00 | 32.258% | 16.386% | Baseline | Not estimated |
| bf16_length_b24_projcache_repstop | 28m 25s | 76.82 | 11m 21s | 192.25 | 31.320% | 14.241% | -0.937 pp | [-1.48, -0.48] pp |
| Approach | Segments | Total | Prepare | Generate | Peak alloc GB | Peak reserved GB |
|---|---|---|---|---|---|---|
| bf16_chron_full | 365 | 64.64s | 12.16s | 52.20s | 6.52 | 11.44 |
| bf16_sorted_full | 365 | 26.61s | 3.98s | 22.55s | 6.47 | 6.95 |
| bf16_rep_cpu_features | 96 | 8.82s | 1.50s | 7.30s | 6.52 | 6.95 |
| bf16_rep_gpu_features | 96 | 8.65s | 1.33s | 7.28s | 6.52 | 6.95 |
| bf16_rep_eager | 96 | 16.08s | 1.30s | 14.73s | 6.52 | 6.95 |
| fp16_sorted_full | 365 | 34.26s | 3.95s | 30.24s | 6.48 | 12.08 |
- Length sorting reduced the 365-segment ASR-stage probe from 64.64s to 26.61s (2.43× faster). Exact strings changed on 14 of 365 segments; full-suite paired WER changed by −0.043 pp, 95% CI [−0.214, +0.081].
- FP16 generation was slower than BF16 on this hardware (34.26s vs 26.61s) and changed 24 segment strings.
- Moving representative features to GPU saved little (8.82s to 8.65s); eager per-item processing was much slower (16.08s).
- Batch 32 consumed more memory and split once without beating batch 24; batch 16 avoided OOM but was slower. Batch 24 length sorting was the pragmatic winner.
Production-pipeline changes
| Area | Measured or functional result | Implementation boundary |
|---|---|---|
| Silero VAD runtime | 5.918s worker time on 1.wav | Vectorized ONNX CPU inference, sample-index boundaries, isolated TorchScript fallback, and provider provenance |
| Feature pipeline | 26.0s to 22.2s in the sorted ASR pass | Prepare one batch ahead while the GPU generates; decoded text was identical in this probe |
| Corpus batching | One ASR and one aligner load per input set | Global stable segment identities and optional duration sorting across files |
| Alignment memory | Roughly 800 MB less redundant allocation on an hour-scale file | Construct CTC windows a batch at a time; input windows and ordered outputs were parity-checked |
| Resource controls | Bounded decoded-audio groups and worker queues | Serialized GPU work, conservative CPU preprocessing, persistent OOM cap learning, and frame-balanced split/retry |
| Decoder completeness | Missing EOS at the token ceiling is no longer silent | Retry only affected rows up to the model positional cap and record unresolved rows prominently |
| Decoder hot path | Projection/mask work cached across autoregressive steps | Model-specific hooks are guarded by the validated Transformers 5.13.x compatibility bound |
| Batch-file behavior | Multiple paths and recursive folders | Collision checks, preserved relative paths, per-file failure isolation, and one output set per source |
| Alignment cleanup | FP16 folder wall 192.24s to 119.07s | Run full Python GC every 64 files instead of after every short clip |
| Alignment numerics | Finite FP32 log probabilities under extreme-logit regression | Crop on accelerator and use torch.log_softmax(logits.float()) before host transfer |
| Viterbi safety | 500/500 files completed after a native signal-11 failure | Use TorchAudio's maintained CTC forced-align operation instead of the package's unsafe ctypes buffer wrapper |
| Merged subtitle timing | Known pauses remain visible | Retain raw VAD speech spans and distribute approximate words over speech rather than intervening silence |
| Decode memory/provenance | FFmpeg streams into one backing buffer | Record the concrete TorchCodec, librosa, or FFmpeg backend and reuse it for alignment |
| Output publication | Transactional TXT/SRT/VTT/JSON writes | Raw indexed words, cues, rollback, temporary cleanup, permission preservation, and source-mutation detection |
| Optimization telemetry | 6.06/6.10 GiB allocated/reserved | Exact per-batch rows, frames, padding, tokens, transfers, generation, OOM, and truncation events in atomic JSON |
Rejected or non-default trials: torch.compile recompiled dynamic shapes and hurt this one-shot workload; FP16 ASR was slower and changed text; eager attention, oversized adaptive batches, pinned host memory, and shorter CTC variants did not provide a reliable measured win. FP16 alignment remains available explicitly, but it can shift word timestamps and is therefore not the parity default. The actual decoder is selected per installed environment and recorded in profile provenance; the final 1.wav gate resolved to FFmpeg.
Native GGUF batch engine
The new cohere-batch executable is a separate C++/GGML research runtime. It loads one GGUF model for every supplied file and directory, plans long-audio windows globally, computes deterministic CPU features in parallel, batches encoder and ragged decoder work, can select greedy tokens on the GPU, and writes collision-safe plain-text outputs atomically. It does not use Silero, forced alignment, or the Python model implementation.
| Model | Internal wall | External wall | Steady RTFx | Load | Encoder | Decoder | Peak VRAM | Lexical WER | Delta vs F16 | Paired 95% CI | Lexical changes |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Native F16 | 109.866s | 110.204s | 46.19 | 0.853s | 46.265s | 57.528s | 7,278 MiB | 22.629% | Reference | Reference | 0 / 500 |
| Native Q8_0 | 76.920s | 77.330s | 65.94 | 0.546s | 39.834s | 31.202s | 5,062 MiB | 22.784% | +0.155 pp | [-0.045, +0.437] pp | 20 / 500 |
| Native Q4_K imatrix | 63.160s | 63.414s | 80.21 | 0.382s | 40.099s | 17.345s | 4,188 MiB | 23.157% | +0.528 pp | [-0.315, +1.609] pp | 153 / 500 |
| Domain | F16 | Q8_0 | Q4_K imatrix |
|---|---|---|---|
| MSA | 6.465% | 6.424% (-0.040 pp; CI [-0.131, +0.000]) | 6.626% (+0.162 pp; CI [-0.165, +0.511]) |
| Dialect | 42.604% | 43.028% (+0.424 pp; CI [-0.043, +1.120]) | 44.299% (+1.695 pp; CI [-0.112, +4.199]) |
| Classical Arabic | 13.994% | 13.994% (+0.000 pp; CI [+0.000, +0.000]) | 12.974% (-1.020 pp; CI [-3.066, +0.134]) |
| Dataset | F16 | Q8_0 | Q4_K imatrix |
|---|---|---|---|
| Common Voice 18 Arabic | 5.859% | 5.664% (-0.195 pp) | 6.445% (+0.586 pp) |
| FLEURS ar-EG | 6.623% | 6.623% (+0.000 pp) | 6.673% (+0.051 pp) |
| SADA22 | 36.849% | 37.329% (+0.479 pp) | 37.671% (+0.822 pp) |
| Casablanca | 50.000% | 50.352% (+0.352 pp) | 52.817% (+2.817 pp) |
| Quran ayah CA proxy | 13.994% | 13.994% (+0.000 pp) | 12.974% (-1.020 pp) |
| Engine/config | External wall | RTFx | Peak GPU | Lexical WER | Boundary |
|---|---|---|---|---|---|
| Python BF16, fixed 30 s, no VAD | 49.58s | 101.57 | 6.07 GiB | 23.266% | HF runtime; plain text; fixed 30 s planner |
| Native F16 | 110.20s | 45.67 | 7.11 GiB | 22.629% | GGUF runtime; plain text; reference quiet-boundary planner |
| Native Q8_0 | 77.33s | 65.08 | 4.94 GiB | 22.784% | GGUF runtime; plain text; reference quiet-boundary planner |
| Native Q4_K imatrix | 63.41s | 79.36 | 4.09 GiB | 23.157% | GGUF runtime; plain text; reference quiet-boundary planner |
| Wit.ai through Tafrigh, 8 apps | 123.50s | 40.78 | Cloud | 34.472% | Auditok, MP3 payloads, cloud requests |
The cross-engine table uses the same 500 source clips and references, but it is not a controlled decoder-only comparison: the Python row uses fixed 30-second no-VAD chunks, native uses the pinned quiet-boundary planner up to 35 seconds, and Tafrigh uses Auditok plus MP3 cloud requests. On this workload the established Python BF16 path remained faster than every native variant. Native Q8 is therefore an implementation and deployment alternative, not a new overall speed record.
- After applying the production repetition cleanup, batched F16 matched corrected serial native F16 on 500/500 raw transcripts.
- On-device argmax preserves greedy decoding but omits confidence logits. Q8
1.wavwas text-identical with and without full-logit readback. - Parallel native feature extraction reduced the fixed-window
1.wavfrontend from 11.822s at one thread to 4.121s at four and 3.347s at twelve; all feature buffers were byte-identical. - Stable encoder frame buckets were slower. A persistent decoder graph was 0.7% faster on a 20-minute Q8 trial but inserted one punctuation token, so both experiments remain non-default.
- A repaired shared-weight CUDA multistream probe was byte-exact, but decoder-like B2 overlap reached only 1.084x and B4 1.147x; the predefined 1.15x B2-first gate failed, so no model/session refactor was attempted.
- The native executable writes text only. It provides neither segment nor word timestamps, so the timestamp-error tables above do not apply to it.
Upstream contribution opportunities
The final script was mapped against current default-branch source in four upstream repositories. GitHub searches covered open and closed issues and pull requests, exact symbols, broader concepts, and nearby changes. “No duplicate found” means no matching public indexed work was found as of July 12, 2026; it cannot exclude an unlinked draft or private branch.
| Rank | Candidate | Target | Evidence | Duplicate check and action |
|---|---|---|---|---|
| 1 · P0 | Generic repeated-block stopping criterion | Transformers generation | Local per-row sticky criterion stopped pathological loops; 500-clip wall 59.603→51.067s after projection cache and WER improved on the full suite | Existing issue #32902; no linked/searchable PR. Comment with ASR evidence and contribute there. |
| 2 · P0 | Project Cohere encoder states once per autoregressive decode | Transformers Cohere ASR | Current decoder repeats an invariant projection per generated token; local ablation 61.208→59.603s wall, 0/500 text changes | Issue search / PR search: no open/closed duplicate. #45214 only handled device placement. |
| 3 · P0 | Stable ONNX emission log-softmax | deskpai/ctc_forced_aligner | Current NumPy log(exp(x)/sum(exp(x))) overflows; extreme-logit regression is non-finite while the shifted/log-softmax form is finite | Issue search / PR search: no duplicate; submit a focused issue and PR. |
| 4 · P1 | Strict Cohere chunk-text reassembly | Transformers Cohere processor | Reject mismatched/duplicate/gapped metadata and handle all-empty chunks instead of silent zip truncation or IndexError | Issue search / PR search: no duplicate; suitable small processor PR. |
| 5 · P1 | Prepare Cohere cross-attention mask once | Transformers Cohere ASR | Current decoder rebuilds an invariant mask per token; local 500-clip generation 28.533→28.171s, 0/500 text changes | Merged #46738 is adjacent infrastructure, not a duplicate; current main still repeats the work. |
| 6 · P1 | State-preserving vectorized offline Silero ONNX | snakers4/silero-vad | Batch sequence frames while carrying recurrent state; exact ONNX/JIT span parity locally and ~5.8s VAD worker time for 4,161s audio | Concept exists in discussion #408 and #217, but no implementation PR. Coordinate before coding. |
| 7 · P1 | Bounded streaming CTC emissions and learned OOM cap | MahmoudAshraf97 + DeskPai aligners | Normalize/crop each batch, copy directly to bounded host storage, preserve completed windows after OOM, and reuse the learned cap | Open issue #89 reports long-file failures; no streaming PR. Reply with a profile/reproducer first. |
<star> vocabulary index; only revive it with a new current-main failure fixture.Recommended contribution order: first the stable DeskPai log-softmax because it is a tiny correctness patch; then the Cohere projection change with a profiler/call-count test; then the generic repetition criterion on the already accepted feature request. Keep the projection and mask patches separate so training, direct decoder calls, beam expansion, compilation, and multi-device behavior can be reviewed independently. For Silero, contribute reproducible export code and exact 8/16 kHz parity tests rather than only the ONNX binary.
The complete audit, pinned upstream commit hashes, exact local line mappings, search queries, neighboring PR analysis, and proposed tests are in UPSTREAM_CANDIDATES_20260712.md.
Wit/Tafrigh configuration probe
The duration-stratified probe contains 500 clips, 100 from each dataset (1.398810 decoded hours). It tests Tafrigh-compatible 15/16-second cuts, its actual generated-noise padding, true silence, and no padding. The baseline is exact Tafrigh v1.7.8 behavior.
| Configuration | Wall | RTFx | Requests | WER | CER | Δ vs default | Paired 95% CI | Exact disagreements |
|---|---|---|---|---|---|---|---|---|
| 15 s, Tafrigh noise (default) | 123.50s | 40.78 | 894 | 34.47% | 24.91% | Baseline | Not estimated | 0 / 0.00% |
| 15 s, true silence | 125.52s | 40.12 | 894 | 34.07% | 23.89% | -0.40 pp | [-2.25, +1.65] pp | 241 / 48.20% |
| 16 s, Tafrigh noise | 123.01s | 40.94 | 880 | 33.93% | 24.57% | -0.54 pp | [-1.23, +0.07] pp | 151 / 30.20% |
| 16 s, true silence | 124.38s | 40.49 | 880 | 33.88% | 23.93% | -0.59 pp | [-2.44, +1.65] pp | 232 / 46.40% |
| 16 s, no padding | 121.57s | 41.42 | 880 | 42.70% | 33.22% | +8.23 pp | [+1.13, +18.40] pp | 354 / 70.80% |
| Configuration | Profile | WER | CER | WER Δ vs default | WER S/D/I |
|---|---|---|---|---|---|
| 15 s, Tafrigh noise (default) | Raw | 57.30% | 36.98% | Baseline | 2,198/1,398/98 |
| 15 s, Tafrigh noise (default) | Leaderboard repo-exact | 37.47% | 25.65% | Baseline | 902/1,407/107 |
| 15 s, Tafrigh noise (default) | Leaderboard intended | 35.01% | 25.08% | Baseline | 745/1,402/107 |
| 15 s, Tafrigh noise (default) | Lexical normalized | 34.47% | 24.91% | Baseline | 714/1,403/104 |
| 15 s, true silence | Raw | 55.64% | 36.07% | -1.66 pp | 2,174/1,308/105 |
| 15 s, true silence | Leaderboard repo-exact | 37.02% | 24.59% | -0.45 pp | 952/1,319/116 |
| 15 s, true silence | Leaderboard intended | 34.57% | 24.05% | -0.43 pp | 797/1,312/117 |
| 15 s, true silence | Lexical normalized | 34.07% | 23.89% | -0.40 pp | 768/1,313/114 |
| 16 s, no padding | Raw | 57.61% | 43.84% | +0.31 pp | 1,667/1,964/83 |
| 16 s, no padding | Leaderboard repo-exact | 45.34% | 33.85% | +7.86 pp | 864/1,970/89 |
| 16 s, no padding | Leaderboard intended | 42.94% | 33.37% | +7.94 pp | 714/1,962/89 |
| 16 s, no padding | Lexical normalized | 42.70% | 33.22% | +8.23 pp | 702/1,963/86 |
| 16 s, Tafrigh noise | Raw | 57.08% | 36.69% | -0.22 pp | 2,201/1,378/101 |
| 16 s, Tafrigh noise | Leaderboard repo-exact | 36.81% | 25.31% | -0.67 pp | 876/1,387/110 |
| 16 s, Tafrigh noise | Leaderboard intended | 34.43% | 24.75% | -0.57 pp | 726/1,381/110 |
| 16 s, Tafrigh noise | Lexical normalized | 33.93% | 24.57% | -0.54 pp | 697/1,382/107 |
| 16 s, true silence | Raw | 55.61% | 36.12% | -1.69 pp | 2,164/1,323/98 |
| 16 s, true silence | Leaderboard repo-exact | 36.76% | 24.63% | -0.71 pp | 929/1,333/108 |
| 16 s, true silence | Leaderboard intended | 34.38% | 24.10% | -0.62 pp | 779/1,326/109 |
| 16 s, true silence | Lexical normalized | 33.88% | 23.93% | -0.59 pp | 750/1,327/106 |
Probe results by dataset and domain
| Dataset | Configuration | WER | CER | Δ | Paired 95% CI | Exact disagreements |
|---|---|---|---|---|---|---|
| Casablanca | 15 s, Tafrigh noise (default) | 50.09% | 30.67% | Baseline | Not estimated | 0 / 0.00% |
| Casablanca | 15 s, true silence | 48.59% | 28.10% | -1.50 pp | [-3.67, +0.47] pp | 60 / 60.00% |
| Casablanca | 16 s, no padding | 49.56% | 28.38% | -0.53 pp | [-2.74, +1.25] pp | 72 / 72.00% |
| Casablanca | 16 s, Tafrigh noise | 49.65% | 30.02% | -0.44 pp | [-1.48, +0.50] pp | 32 / 32.00% |
| Casablanca | 16 s, true silence | 48.42% | 27.77% | -1.67 pp | [-3.41, -0.10] pp | 57 / 57.00% |
| Common Voice 18 Arabic | 15 s, Tafrigh noise (default) | 12.30% | 6.09% | Baseline | Not estimated | 0 / 0.00% |
| Common Voice 18 Arabic | 15 s, true silence | 9.57% | 3.02% | -2.73 pp | [-5.43, -0.20] pp | 24 / 24.00% |
| Common Voice 18 Arabic | 16 s, no padding | 19.53% | 12.37% | +7.23 pp | [+1.82, +13.00] pp | 55 / 55.00% |
| Common Voice 18 Arabic | 16 s, Tafrigh noise | 11.52% | 5.21% | -0.78 pp | [-2.52, +0.89] pp | 16 / 16.00% |
| Common Voice 18 Arabic | 16 s, true silence | 9.57% | 3.26% | -2.73 pp | [-5.15, -0.55] pp | 21 / 21.00% |
| FLEURS ar-EG | 15 s, Tafrigh noise (default) | 26.44% | 20.94% | Baseline | Not estimated | 0 / 0.00% |
| FLEURS ar-EG | 15 s, true silence | 24.50% | 18.57% | -1.94 pp | [-3.04, -0.96] pp | 39 / 39.00% |
| FLEURS ar-EG | 16 s, no padding | 28.32% | 22.51% | +1.88 pp | [+0.39, +3.40] pp | 80 / 80.00% |
| FLEURS ar-EG | 16 s, Tafrigh noise | 26.44% | 20.93% | +0.00 pp | [-0.68, +0.71] pp | 26 / 26.00% |
| FLEURS ar-EG | 16 s, true silence | 24.30% | 18.44% | -2.14 pp | [-3.46, -1.09] pp | 40 / 40.00% |
| Quran ayah CA proxy | 15 s, Tafrigh noise (default) | 39.80% | 36.13% | Baseline | Not estimated | 0 / 0.00% |
| Quran ayah CA proxy | 15 s, true silence | 45.26% | 41.16% | +5.47 pp | [+0.21, +11.29] pp | 56 / 56.00% |
| Quran ayah CA proxy | 16 s, no padding | 73.32% | 71.75% | +33.53 pp | [+7.90, +70.43] pp | 79 / 79.00% |
| Quran ayah CA proxy | 16 s, Tafrigh noise | 37.97% | 35.17% | -1.82 pp | [-5.47, -0.21] pp | 37 / 37.00% |
| Quran ayah CA proxy | 16 s, true silence | 44.46% | 40.80% | +4.66 pp | [-0.91, +12.54] pp | 60 / 60.00% |
| SADA22 | 15 s, Tafrigh noise (default) | 35.89% | 22.07% | Baseline | Not estimated | 0 / 0.00% |
| SADA22 | 15 s, true silence | 33.70% | 18.95% | -2.19 pp | [-3.80, -0.71] pp | 62 / 62.00% |
| SADA22 | 16 s, no padding | 36.03% | 22.07% | +0.14 pp | [-1.88, +2.27] pp | 68 / 68.00% |
| SADA22 | 16 s, Tafrigh noise | 35.82% | 22.28% | -0.07 pp | [-1.14, +1.04] pp | 40 / 40.00% |
| SADA22 | 16 s, true silence | 34.04% | 19.90% | -1.85 pp | [-3.70, -0.26] pp | 54 / 54.00% |
| Domain | Configuration | WER | CER | Δ | Paired 95% CI | Exact disagreements |
|---|---|---|---|---|---|---|
| Classical Arabic | 15 s, Tafrigh noise (default) | 39.80% | 36.13% | Baseline | Not estimated | 0 / 0.00% |
| Classical Arabic | 15 s, true silence | 45.26% | 41.16% | +5.47 pp | [+0.21, +11.29] pp | 56 / 56.00% |
| Classical Arabic | 16 s, no padding | 73.32% | 71.75% | +33.53 pp | [+7.90, +70.43] pp | 79 / 79.00% |
| Classical Arabic | 16 s, Tafrigh noise | 37.97% | 35.17% | -1.82 pp | [-5.47, -0.21] pp | 37 / 37.00% |
| Classical Arabic | 16 s, true silence | 44.46% | 40.80% | +4.66 pp | [-0.91, +12.54] pp | 60 / 60.00% |
| Dialect | 15 s, Tafrigh noise (default) | 42.10% | 25.76% | Baseline | Not estimated | 0 / 0.00% |
| Dialect | 15 s, true silence | 40.22% | 22.88% | -1.89 pp | [-3.16, -0.70] pp | 122 / 61.00% |
| Dialect | 16 s, no padding | 41.95% | 24.78% | -0.15 pp | [-1.57, +1.28] pp | 140 / 70.00% |
| Dialect | 16 s, Tafrigh noise | 41.87% | 25.60% | -0.23 pp | [-0.97, +0.54] pp | 72 / 36.00% |
| Dialect | 16 s, true silence | 40.33% | 23.28% | -1.77 pp | [-3.08, -0.64] pp | 111 / 55.50% |
| MSA | 15 s, Tafrigh noise (default) | 23.52% | 18.20% | Baseline | Not estimated | 0 / 0.00% |
| MSA | 15 s, true silence | 21.41% | 15.70% | -2.10 pp | [-3.10, -1.13] pp | 63 / 31.50% |
| MSA | 16 s, no padding | 26.51% | 20.63% | +2.99 pp | [+1.38, +4.73] pp | 135 / 67.50% |
| MSA | 16 s, Tafrigh noise | 23.35% | 18.03% | -0.16 pp | [-0.79, +0.45] pp | 42 / 21.00% |
| MSA | 16 s, true silence | 21.25% | 15.64% | -2.26 pp | [-3.33, -1.22] pp | 61 / 30.50% |
Probe results for every dialect label
| Dialect | Configuration | WER | CER | Δ | Paired 95% CI | Exact disagreements |
|---|---|---|---|---|---|---|
| Algeria | 15 s, Tafrigh noise (default) | 66.07% | 33.57% | Baseline | Not estimated | 0 / 0.00% |
| Algeria | 15 s, true silence | 59.82% | 29.37% | -6.25 pp | [-12.87, +0.00] pp | 7 / 50.00% |
| Algeria | 16 s, no padding | 63.39% | 28.67% | -2.68 pp | [-10.59, +3.18] pp | 10 / 71.43% |
| Algeria | 16 s, Tafrigh noise | 64.29% | 31.47% | -1.79 pp | [-6.67, +2.22] pp | 4 / 28.57% |
| Algeria | 16 s, true silence | 60.71% | 29.90% | -5.36 pp | [-12.82, +1.96] pp | 9 / 64.29% |
| Classical Arabic | 15 s, Tafrigh noise (default) | 39.80% | 36.13% | Baseline | Not estimated | 0 / 0.00% |
| Classical Arabic | 15 s, true silence | 45.26% | 41.16% | +5.47 pp | [+0.21, +11.29] pp | 56 / 56.00% |
| Classical Arabic | 16 s, no padding | 73.32% | 71.75% | +33.53 pp | [+7.90, +70.43] pp | 79 / 79.00% |
| Classical Arabic | 16 s, Tafrigh noise | 37.97% | 35.17% | -1.82 pp | [-5.47, -0.21] pp | 37 / 37.00% |
| Classical Arabic | 16 s, true silence | 44.46% | 40.80% | +4.66 pp | [-0.91, +12.54] pp | 60 / 60.00% |
| Egypt | 15 s, Tafrigh noise (default) | 30.58% | 12.55% | Baseline | Not estimated | 0 / 0.00% |
| Egypt | 15 s, true silence | 32.52% | 14.20% | +1.94 pp | [-0.96, +6.17] pp | 9 / 60.00% |
| Egypt | 16 s, no padding | 33.50% | 13.17% | +2.91 pp | [-2.08, +7.78] pp | 12 / 80.00% |
| Egypt | 16 s, Tafrigh noise | 30.10% | 12.04% | -0.49 pp | [-1.59, +0.00] pp | 4 / 26.67% |
| Egypt | 16 s, true silence | 31.55% | 11.93% | +0.97 pp | [-1.40, +3.29] pp | 8 / 53.33% |
| Egyptian | 15 s, Tafrigh noise (default) | 25.00% | 5.56% | Baseline | Not estimated | 0 / 0.00% |
| Egyptian | 15 s, true silence | 25.00% | 8.33% | +0.00 pp | [+0.00, +0.00] pp | 1 / 50.00% |
| Egyptian | 16 s, no padding | 25.00% | 9.26% | +0.00 pp | [-7.69, +14.29] pp | 2 / 100.00% |
| Egyptian | 16 s, Tafrigh noise | 25.00% | 8.33% | +0.00 pp | [+0.00, +0.00] pp | 1 / 50.00% |
| Egyptian | 16 s, true silence | 25.00% | 8.33% | +0.00 pp | [+0.00, +0.00] pp | 1 / 50.00% |
| Hijazi | 15 s, Tafrigh noise (default) | 29.85% | 22.60% | Baseline | Not estimated | 0 / 0.00% |
| Hijazi | 15 s, true silence | 21.64% | 13.00% | -8.21 pp | [-15.18, -3.23] pp | 10 / 58.82% |
| Hijazi | 16 s, no padding | 29.85% | 21.83% | +0.00 pp | [-4.35, +4.40] pp | 11 / 64.71% |
| Hijazi | 16 s, Tafrigh noise | 29.10% | 20.74% | -0.75 pp | [-4.40, +2.91] pp | 6 / 35.29% |
| Hijazi | 16 s, true silence | 23.88% | 13.62% | -5.97 pp | [-13.27, -0.73] pp | 8 / 47.06% |
| Jordan | 15 s, Tafrigh noise (default) | 29.27% | 9.47% | Baseline | Not estimated | 0 / 0.00% |
| Jordan | 15 s, true silence | 30.49% | 8.16% | +1.22 pp | [-8.64, +14.93] pp | 6 / 75.00% |
| Jordan | 16 s, no padding | 28.05% | 8.42% | -1.22 pp | [-6.82, +4.82] pp | 4 / 50.00% |
| Jordan | 16 s, Tafrigh noise | 30.49% | 10.53% | +1.22 pp | [+0.00, +3.57] pp | 4 / 50.00% |
| Jordan | 16 s, true silence | 28.05% | 7.37% | -1.22 pp | [-9.38, +9.59] pp | 6 / 75.00% |
| Khaliji | 15 s, Tafrigh noise (default) | 52.17% | 37.43% | Baseline | Not estimated | 0 / 0.00% |
| Khaliji | 15 s, true silence | 46.09% | 30.20% | -6.09 pp | [-15.46, +0.49] pp | 11 / 57.89% |
| Khaliji | 16 s, no padding | 51.30% | 38.88% | -0.87 pp | [-7.50, +4.17] pp | 10 / 52.63% |
| Khaliji | 16 s, Tafrigh noise | 54.78% | 41.59% | +2.61 pp | [-1.77, +8.18] pp | 5 / 26.32% |
| Khaliji | 16 s, true silence | 49.57% | 33.82% | -2.61 pp | [-8.97, +3.38] pp | 8 / 42.11% |
| MSA | 15 s, Tafrigh noise (default) | 23.52% | 18.20% | Baseline | Not estimated | 0 / 0.00% |
| MSA | 15 s, true silence | 21.41% | 15.70% | -2.10 pp | [-3.05, -1.21] pp | 63 / 31.50% |
| MSA | 16 s, no padding | 26.51% | 20.63% | +2.99 pp | [+1.34, +4.61] pp | 135 / 67.50% |
| MSA | 16 s, Tafrigh noise | 23.35% | 18.03% | -0.16 pp | [-0.83, +0.48] pp | 42 / 21.00% |
| MSA | 16 s, true silence | 21.25% | 15.64% | -2.26 pp | [-3.37, -1.29] pp | 61 / 30.50% |
| Mauritania | 15 s, Tafrigh noise (default) | 89.95% | 83.41% | Baseline | Not estimated | 0 / 0.00% |
| Mauritania | 15 s, true silence | 92.06% | 85.70% | +2.12 pp | [+0.00, +5.14] pp | 5 / 38.46% |
| Mauritania | 16 s, no padding | 87.83% | 79.98% | -2.12 pp | [-8.29, +1.82] pp | 7 / 53.85% |
| Mauritania | 16 s, Tafrigh noise | 88.89% | 81.58% | -1.06 pp | [-3.68, +1.23] pp | 3 / 23.08% |
| Mauritania | 16 s, true silence | 88.89% | 83.41% | -1.06 pp | [-4.80, +1.58] pp | 4 / 30.77% |
| More than 1 speaker اكثر من متحدث | 15 s, Tafrigh noise (default) | 35.12% | 21.35% | Baseline | Not estimated | 0 / 0.00% |
| More than 1 speaker اكثر من متحدث | 15 s, true silence | 34.33% | 19.88% | -0.78 pp | [-3.05, +1.18] pp | 16 / 76.19% |
| More than 1 speaker اكثر من متحدث | 16 s, no padding | 34.60% | 20.44% | -0.52 pp | [-3.44, +2.62] pp | 19 / 90.48% |
| More than 1 speaker اكثر من متحدث | 16 s, Tafrigh noise | 34.99% | 21.86% | -0.13 pp | [-1.80, +1.66] pp | 17 / 80.95% |
| More than 1 speaker اكثر من متحدث | 16 s, true silence | 34.07% | 21.14% | -1.04 pp | [-3.64, +1.14] pp | 15 / 71.43% |
| Morocco | 15 s, Tafrigh noise (default) | 47.49% | 23.43% | Baseline | Not estimated | 0 / 0.00% |
| Morocco | 15 s, true silence | 39.66% | 14.08% | -7.82 pp | [-16.32, -0.56] pp | 12 / 70.59% |
| Morocco | 16 s, no padding | 45.25% | 16.45% | -2.23 pp | [-10.06, +3.43] pp | 13 / 76.47% |
| Morocco | 16 s, Tafrigh noise | 48.60% | 23.91% | +1.12 pp | [-1.90, +3.89] pp | 6 / 35.29% |
| Morocco | 16 s, true silence | 45.25% | 19.64% | -2.23 pp | [-7.65, +1.96] pp | 9 / 52.94% |
| Najdi | 15 s, Tafrigh noise (default) | 26.57% | 12.48% | Baseline | Not estimated | 0 / 0.00% |
| Najdi | 15 s, true silence | 25.37% | 11.02% | -1.19 pp | [-3.50, +0.73] pp | 17 / 77.27% |
| Najdi | 16 s, no padding | 26.87% | 12.61% | +0.30 pp | [-2.49, +3.17] pp | 17 / 77.27% |
| Najdi | 16 s, Tafrigh noise | 26.57% | 12.24% | +0.00 pp | [-1.11, +0.87] pp | 8 / 36.36% |
| Najdi | 16 s, true silence | 26.27% | 11.02% | -0.30 pp | [-2.50, +1.56] pp | 14 / 63.64% |
| Notapplicable | 15 s, Tafrigh noise (default) | 70.00% | 57.45% | Baseline | Not estimated | 0 / 0.00% |
| Notapplicable | 15 s, true silence | 60.00% | 34.04% | -10.00 pp | [-20.00, +0.00] pp | 1 / 33.33% |
| Notapplicable | 16 s, no padding | 70.00% | 57.45% | +0.00 pp | [+0.00, +0.00] pp | 0 / 0.00% |
| Notapplicable | 16 s, Tafrigh noise | 50.00% | 46.81% | -20.00 pp | [-100.00, +0.00] pp | 1 / 33.33% |
| Notapplicable | 16 s, true silence | 60.00% | 34.04% | -10.00 pp | [-20.00, +0.00] pp | 1 / 33.33% |
| Palestine | 15 s, Tafrigh noise (default) | 28.36% | 10.03% | Baseline | Not estimated | 0 / 0.00% |
| Palestine | 15 s, true silence | 34.33% | 10.64% | +5.97 pp | [+1.20, +13.64] pp | 5 / 50.00% |
| Palestine | 16 s, no padding | 34.33% | 16.72% | +5.97 pp | [+0.00, +12.68] pp | 6 / 60.00% |
| Palestine | 16 s, Tafrigh noise | 29.85% | 9.42% | +1.49 pp | [+0.00, +6.52] pp | 4 / 40.00% |
| Palestine | 16 s, true silence | 29.85% | 9.42% | +1.49 pp | [+0.00, +6.52] pp | 4 / 40.00% |
| UAE | 15 s, Tafrigh noise (default) | 50.00% | 29.55% | Baseline | Not estimated | 0 / 0.00% |
| UAE | 15 s, true silence | 44.34% | 22.11% | -5.66 pp | [-11.24, +0.00] pp | 10 / 83.33% |
| UAE | 16 s, no padding | 46.23% | 23.97% | -3.77 pp | [-7.29, +0.00] pp | 11 / 91.67% |
| UAE | 16 s, Tafrigh noise | 50.00% | 30.37% | +0.00 pp | [+0.00, +0.00] pp | 2 / 16.67% |
| UAE | 16 s, true silence | 45.28% | 22.31% | -4.72 pp | [-10.23, +0.00] pp | 9 / 75.00% |
| Unknown | 15 s, Tafrigh noise (default) | 67.50% | 46.82% | Baseline | Not estimated | 0 / 0.00% |
| Unknown | 15 s, true silence | 63.75% | 38.17% | -3.75 pp | [-8.97, +1.64] pp | 6 / 37.50% |
| Unknown | 16 s, no padding | 75.00% | 53.18% | +7.50 pp | [-4.17, +20.51] pp | 9 / 56.25% |
| Unknown | 16 s, Tafrigh noise | 67.50% | 44.53% | +0.00 pp | [+0.00, +0.00] pp | 2 / 12.50% |
| Unknown | 16 s, true silence | 60.00% | 37.40% | -7.50 pp | [-15.19, +0.00] pp | 7 / 43.75% |
| Yemen | 15 s, Tafrigh noise (default) | 41.54% | 21.26% | Baseline | Not estimated | 0 / 0.00% |
| Yemen | 15 s, true silence | 40.00% | 17.78% | -1.54 pp | [-6.71, +2.72] pp | 6 / 54.55% |
| Yemen | 16 s, no padding | 41.54% | 20.94% | +0.00 pp | [-3.79, +2.30] pp | 9 / 81.82% |
| Yemen | 16 s, Tafrigh noise | 39.49% | 19.96% | -2.05 pp | [-5.62, +0.00] pp | 5 / 45.45% |
| Yemen | 16 s, true silence | 39.49% | 15.59% | -2.05 pp | [-7.48, +1.46] pp | 8 / 72.73% |
Tafrigh calls WhiteNoise(..., volume=0). In Pydub, 0 means 0 dB gain rather than zero amplitude; measured output was roughly −4.8 dBFS. “True silence” used −120 dB. Wit outputs were not deterministic across calls/apps, including one minor Quran variant on the same encoded payload.
Wit/Tafrigh pipeline and telemetry
| Measure | Value | Meaning |
|---|---|---|
| Successful speech-segment requests | 32,160 | HTTP 200 segment recognitions |
| Recorded HTTP attempts | 32,192 | Includes retries |
| Recovered retries | 32 | 0.099% of attempts |
| Unresolved failures | 0 | No transport/API failure scored as empty text |
| No-speech clips | 82 | Auditok detected no region |
| Successful empty segments | 4,980 | HTTP 200 with empty text |
| Clips with one or more empty segments | 3,954 | Still checkpointed as successful |
| Detected speech | 31.731 h | 87.19% of decoded input |
| Whole-source WAV→MP3 conversions | 13,943 | Tafrigh compatibility preprocessing |
| MP3 request payload | 1.793 GB | Aggregate uploaded bytes |
| Forced MP3 demux fallbacks | 1 | One valid file was mis-probed by Auditok/FFmpeg |
| Preprocessing retries | 0 | None required |
Compatibility boundary
- Non-MP3 clips are converted in full to MP3 with Pydub/FFmpeg.
- Tafrigh's Auditok splitter uses
min_dur=0.5,max_silence=0.5,energy_threshold=50, and a 15-second maximum region. - Each region is re-encoded after one second of Tafrigh-generated noise is added at both sides.
- Requests use
POST /speech,audio/mpeg3, Wit media version20200513, Arabic question-mark substitution, and chronological segment concatenation.
Scheduler
Stock Tafrigh loops over files sequentially and creates a manager/process pool for every file. On 24,414 short clips, even a one-request-per-file admission lower bound is 6h 47m. The benchmark preserves Tafrigh's recognition inputs but shares a bounded scheduler across files, uses 8 distinct Arabic apps in 2 inferred user groups, limits each app to one start every 1.02s, caps each user group at 4 starts/s, and checkpoints a clip only after every segment succeeds. Every retry reacquires both quota limits.
The operator's production topology is eight Tafrigh processes with eight keys each, and the 64 keys have independent app/user quotas. That can materially outperform a single eight-app run on a folder when enough independent files exist: the processes can upload and recognize separate regions concurrently. It does not make one unsharded file automatically eight times faster, and actual aggregate scaling still depends on Auditok/encoding CPU, upload bandwidth, network latency, API service time, and per-process scheduling. Cohere has the opposite shape: one GPU model is best kept resident and globally batched across files rather than duplicated eight times on a 12 GB card.
Methodology and reproducibility
| Item | Specification |
|---|---|
| Primary engines | Cohere Arabic Transcribe model; Wit.ai /speech through Tafrigh 1.7.8-compatible preprocessing |
| Cohere revision | 0a8193caa4f3f92131471ab08824e488141cb392 |
| July 12 production script | d1dda35f9d393ee11001d4fe71c7c6db255f76c4320f90baf354429d7423b510 |
| Production runtime bound | Transformers >=5.13,<5.14; portable dependencies in final/requirements.txt |
| Production GPU | NVIDIA GeForce RTX 3060, 11.754 GiB, compute capability 8.6 |
| Tafrigh revision | 2ccba42db8c34c04924d1befc35cb3d5eec80d93 |
| Native runtime | CrispASR/GGML CUDA research branch; GGML 0714117daca2471b00e09554c7eaa74a06b0b2c5 |
| Native GGUF gate | 500 clips / 5,032.699 s; E4/D8; greedy device argmax; F16, Q8_0, Q4_K imatrix |
| Sample pairing | 24,414 matching IDs, references, metadata records, and audio paths |
| Timing denominator | 131,015.81575 decoded seconds at 16 kHz; 36.393282 hours |
| Primary statistic | Corpus WER: total S+D+I divided by total reference words |
| Primary profile | Lexical normalized |
| Uncertainty | 5,000-replicate paired cluster percentile bootstrap; 95% interval; seed 0 |
| Cluster rule | Speaker within dataset where present; otherwise utterance |
| Overall clusters | 14,319 (976 speaker clusters + 13,343 utterance clusters) |
| Probe uncertainty | 2,000 paired cluster-bootstrap replicates; same confidence and seed |
| Native quantization uncertainty | 20,000 paired utterance-bootstrap replicates; 95% interval; seed 20260711 |
| WER tie-breaking | NeMo/Kaldi-compatible insertion, then deletion, then substitution on equal-cost paths |
Validation gates
- All audio paths decoded successfully; there were 24,414 unique paths.
- IDs, references, dataset/domain/dialect/speaker metadata, and paths matched across configurations.
- Hypothesis-row fingerprints matched their summaries.
- Declared decoder provenance matched where present; the Wit runner did not declare local decoder provenance, so exact duration was independently replayed.
- Wit failures were explicit and resumable; exhausted requests were never silently converted to empty hypotheses.
- Credentials were accepted only through a hidden prompt and were not written to reports, summaries, hypotheses, or logs.
- The July 12 release passed repeated 1.wav timing, three current 500-clip production modes, real FP32 word alignment, ONNX/JIT VAD parity, live truncation retry, 71 automated benchmark tests, Ruff, and bytecode compilation.
Environment
Wit used Python 3.12.12, tafrigh[wit]==1.7.8, and Requests 2.34.2 in an isolated environment without PyTorch, Transformers, Whisper, or Cohere dependencies. The final production profile used Python 3.12.12, PyTorch/TorchAudio 2.11.0+cu128, Transformers 5.13.0, ONNX Runtime 1.27.0, and CUDA BF16 on an NVIDIA GeForce RTX 3060. Exact tested package versions remain recorded in the production profile and release summary; the portable bundle intentionally constrains only compatibility-sensitive packages and leaves device-specific Torch wheels to the installer.
Limitations and correct interpretation
- The 24,414-clip recognition benchmark is pre-segmented and does not measure Silero long-form boundary quality. The separate 500-clip production probe includes Silero, subtitle cueing, and forced alignment. Wit applies Tafrigh/Auditok VAD inside each clip, so its WER includes that package behavior.
- Wit is a cloud service and showed output variation across calls/apps. Cohere timing depends on the tested local GPU and software stack. Each engine has one primary timing observation, so speed has no confidence interval.
- The Wit wall was accumulated across resumed invocations. One component was recovered from progress and checkpoint timestamps to roughly one-second precision; the uncertainty is negligible relative to 4,397 seconds but should still be disclosed.
- The eight-app result is workload-specific. Quota scheduling, network latency, clip count, and average clip duration mean the 2.578× ratio should not be extrapolated directly to one continuous recording.
- The user's 64-key/eight-process Tafrigh topology was not benchmarked as one controlled run. Independent quotas make substantial folder-level scaling plausible, but a linear eightfold extrapolation from the eight-app result would ignore preprocessing, bandwidth, service latency, file distribution, and retry overhead.
- The Quran set has 600 selected ayat but only three reciters, producing a wide cluster-bootstrap interval. Canonical verse text is not an independent verbatim transcript, so this is a Classical Arabic recitation proxy only.
- MSA denotes reference register, not guaranteed speaker accent. Dialect labels combine Casablanca country varieties with SADA annotations and include unknown/multi-speaker categories.
- Very small labels such as Yemeni (7 clips) and Shamali (18 clips) have unstable rates and broad intervals. CER can exceed 100% when insertions outnumber reference characters.
- The lexical normalizer is the closest disclosed approximation to Cohere's internal evaluation normalizer, not a verified byte-identical copy. The repo-exact leaderboard profile intentionally preserves a known punctuation-regex bug.
- None of the five evaluation datasets contains human word-level timestamps. FP32-relative drift measures implementation agreement only; it is not absolute timestamp accuracy and does not prove that FP32 chose the acoustically correct boundary.
- The native GGUF runtime is text-only and uses a different long-audio planner from both the Python fixed-window and Silero paths. Its 500-clip cross-engine table shares sources and references but is not a controlled decoder-only comparison.
- Native quantization intervals resample utterances, not speaker or corpus clusters. Both aggregate intervals include zero; the Q8 recommendation also relies on its much smaller output-change count and domain deltas.
- Three native F16 lanes and three Q8 lanes reached the 445-token generation cap in the 500-clip gate; Q4 had one. The CLI warns and publishes those texts, so maximum-token warnings require review.
- Cloud processing changes the data boundary. Wit states that voice data and transcripts may be retained for up to 90 days. Several evaluation datasets have noncommercial or no-derivatives terms; do not redistribute their audio.
- WER measures transcript edit distance, not semantic fidelity, punctuation quality, named-entity accuracy, timestamp quality, hallucination severity, or downstream task utility.
- The July 12 production patch was gated on the balanced 500 clips rather than rerunning the 24,414-clip suite. It improves production mechanics and sample-exact VAD behavior, but subgroup movement on the 500 clips was not uniform.
- The 36.485-second 1.wav figure is the median of two external process-wall runs on one RTX 3060 system. Model cache state, storage, CPU, driver, thermal state, other GPU users, container decoding, and segment distribution can change results on another device.
- Adaptive growth is memory-aware, not a complete throughput search. Its equal word count but 12-token difference on 1.wav demonstrates why finite-precision configuration changes must retain transcript-quality gates.
Sources, versions, and artifact integrity
Public references
- Cohere Arabic Transcribe technical release and normalization disclosure
- Official Cohere Arabic GGUF variants and quantization notes
- Pinned Transformers Cohere ASR implementation
- TorchAudio maintenance transition update
- CrispASR native runtime source
- Silero VAD and ONNX runtime guidance
- TorchAudio CTC forced-alignment API
- Tafrigh v1.7.8 source
- Wit.ai HTTP /speech documentation
- Wit.ai FAQ
- Wit.ai privacy policy
- Synchronized Cohere/Wit 1.wav playback comparison
- 1.wav source video
- 1.wav hosted audio fallback
- Pinned Arabic ASR leaderboard evaluator and manifests
- Common Voice 18 Arabic mirror
- FLEURS
- SADA22
- Casablanca
- Casablanca corpus paper and utterance-boundary method
- Quran Ayah Corpus used for the CA proxy
Machine-readable source artifacts
| Artifact | SHA-256 | Size/scope |
|---|---|---|
| benchmark/reports/final_cohere_vs_wit_20260711.json | c8306bcd7f70f006e271950a0cf4a052741cfed5e316ffb0a90ebc401f91a12f | 187,895 bytes |
| benchmark/reports/wit_probe500.json | bd73f5668a3e7874038f00dd83949e9e3d3cd6a0190db92cba37425c767c67ab | 439,824 bytes |
| benchmark/reports/full_exact_20260710.json | ea39c3ea18b8c1527c37d03fcebbb71b5aa451fd79e633ae8a3d422b943253d8 | 385,534 bytes |
| benchmark/reports/final_full_20260711.json | bd49322844748b87d34679ac716eaefa4c0694db637b631c0d40253657ad66be | 186,000 bytes |
| benchmark/reports/pipeline_probe500_20260711.json | d77464eaebc2d8b22a6471e77b081d333a53756cf5688664fd67a0c622f8fdda | 168,734 bytes |
| benchmark/results/wit_full_default_20260710/summary.json | 2726e15b57156298849b5315d9fff6bbca471a1983924afda653e1209d3f3c89 | 4,380 bytes |
| benchmark/results/longform_precision_probe.json | 0b7617a54e769009b7ea7b92242640bf8c6e01335dcbd66c98b35298b31216e7 | 4,025 bytes |
| benchmark/reports/timestamp_modes_500_20260711.json | 8947fb3ae5d8023d73114a66e88586143262e58eed773a6e00ce75394db094d1 | 444,127 bytes |
| benchmark/results/timestamp_modes_20260711/performance.json | 590b6cc16c9e49f1b13da1dd8089b9fd065d46fd56fb7e1a96a0b205963d0bd4 | 3,375 bytes |
| benchmark/results/alignment_modes_frozen_20260711/summary.json | ddab3281d0e0d552daef924a2970fa535ee54b30126a9bd8baff97a2f7679947 | 14,855 bytes |
| benchmark/results/longform_timestamp_frozen_20260711/drift.json | 479dcbdab2d5ed3e3ef769f2c2a71752b2aee2fe73cf6a1d4e58e7903f30aede | 149,929 bytes |
| benchmark/results/longform_text_only_20260711/summary.json | c5b883312e65be4f1448a06fcd4fb555d1bf6b1f77f3ea0041ae0537f54eb832 | 692 bytes |
| benchmark/reports/vad_modes_500_20260711.json | 1c1e825dd1ca688a4badbc6d7efb5bef10f8e3e10de128ebe28427bb31220237 | 30,682 bytes |
| benchmark/reports/vad_modes_500_20260711.md | c1ecf1f7938ba42778f362811adc744aa25d61d3a7afcc1b8611b3ce60c50ed4 | 2,658 bytes |
| benchmark/results/vad_modes_20260711/segmentation_audit.json | b587656a53a811218e69fff6364486069f4a9d6f8a27af1fad6d99cb8b32fbeb | 4,241 bytes |
| benchmark/results/vad_modes_20260711/segmentation_audit.time.txt | 50b9c9caf53be327ecbb80133ccb6f9418a4e97dcf0c0f9140c95a049fddde8e | 60 bytes |
| benchmark/reports/production_release_20260712.json | 74f4caf4f49691357ab160711cf07b811d959b7443c2aa470e50d29768e87e2a | 34,982 bytes |
| benchmark/PRODUCTION_REVIEW_20260712.md | 00140f207a86b06c01b2630b33896ed9e0915a9f99df34512f36ed4311f43976 | 9,584 bytes |
| benchmark/reports/cohere_wit_playback_comparison_20260712.html | c0b5250163dae64345afb21ca725b893f8629b17e892d9b5b615edee833f62e3 | 233,834 bytes |
| benchmark/reports/cohere_wit_playback_comparison_20260712.metadata.json | fcd664aaaf0269fbfcb5d677b085aca3682e66b2dd8ba171d97c7ece9f332f16 | 5,172 bytes |
| benchmark/reports/cohere_wit_playback_comparison_20260712.pastehtml.json | ed841063514d088824202352dfe9833c02a7961169a7b2e142f38857e5b147e7 | 654 bytes |
| benchmark/reports/build_cohere_wit_playback_demo.py | edc3b73a1cbd32ee7b11bfef15bac5e3b2ae08b8bac4de52755f96cac089b28b | 34,993 bytes |
| benchmark/reports/cohere_wit_comparison_assets/1-original.json | 99fa81eed7469a06c03c1aa4bb5d6b72c1f80228c1426c5bf818d69dd6816055 | 117,998 bytes |
| benchmark/reports/cohere_wit_comparison_assets/1.json | 28dac6cb48cf4c3607e0be388bebee2fbd347c49880a4616253514c8785d0d49 | 111,880 bytes |
| native_benchmark/results/quantized_full500_20260711/comparison.json | 2cba3089d76024fd469203f2765a7ad431ac63faf1ec55fc41201727f8b7563f | 417,774 bytes |
| native_benchmark/results/quantized_full500_20260711/REPORT.md | 2e52d1c589c9de2268ef5558fbb448b6c43be284bbd96e1d50baf4488b7eb2a4 | 5,786 bytes |
| native_benchmark/results/frontend_parallel/benchmark.json | 17a5a5b4a5b9999ec5022e2dc582be2daeeaa81c30066ef8ff92d6818f8bbc29 | 2,124 bytes |
| native_benchmark/DECODER_BATCH_PARITY.md | 0980c98a065a0382c16a7c1073c02b3d65c03f569ec17381202db00e2a9331e8 | 10,417 bytes |
| native_benchmark/results/multistream_reprobe_20260711/summary.json | 3e6f632c0b4aae989611e32e9852e7f1c9ed7672f0e45991fda00a1770b55ffe | 1,609 bytes |
| native_benchmark/results/multistream_reprobe_20260711/REPORT.md | 86a0136352cebb8c6d7b16cbc9e0ea4bc92a2fdcb1cb04c2cf12c3121bbd2a66 | 4,706 bytes |
| native_benchmark/results/native_onewav_release_20260711/summary.json | 1c28f4e6f7aacba7075e419b088692e3a28cb99717b6b6fa7ed112ccb6e5a27f | 23,362 bytes |
| native_benchmark/results/native_onewav_release_20260711/REPORT.md | b0b75971004f549ca4a1531db31a4d3b5bfdcfaedc0d0fc200a70a7cd3b5b7c3 | 1,911 bytes |
| benchmark/UPSTREAM_CANDIDATES_20260712.md | d84275834d55bfa4a0a6826b92f8c6d7e13f5af356b8e01059da6bcd34c975f3 | 27,209 bytes |
| final/transcribe.py | e27add42e280781e0a98efa1e032a2675fcd9feb459d5e8a2e736e01e1d34e78 | 164,834 bytes |
| final/README.md | 82983546fa37028bb5a95a398d27d34c8809dde509a61f33451658c98ac649f9 | 1,864 bytes |
| final/VERSION.json | 8f4a6215c32133deab8dc5cb015c99e46fd459c38029405cef956191dbff4d96 | 921 bytes |
| final/requirements.txt | 19ecd88291e7bf3c70f43b3761a760470c0aedf2b0a7e9b61cc97d617349ebfb | 982 bytes |
| final/requirements-optional.txt | 33680ee16a154dd87de2c52f06082e7aa7c9472f8d8aa67e9f200063eb7ba292 | 345 bytes |
| final/SETUP.md | 90b78c77fef726aa08132738874889acfd19bb8a617062c7af6f7d43a75fa2a9 | 10,089 bytes |
| final/validate_install.py | a521b02be15354b3f479b0f4889ed85c76994db6edf9e0d6946de1fe2ed71699 | 10,719 bytes |
| final/transcribe_assets/README.md | 3003b2763c12abd0f1d497cded37bf91c6d9094bb6961ff25b15662e17bf2509 | 1,216 bytes |
| final/transcribe_assets/LICENSE.silero-vad | 51c19c8be941a3fb00ccf58f0bf9053de9f7237a0b37327896eabad32dffe873 | 1,076 bytes |
| final/transcribe_assets/LICENSE.faster-whisper | af6798135e729f8aa6c853936d037dfdea449734d26b8ea6a89805fca758c0d5 | 1,064 bytes |
| benchmark/manifests/wer_eval.jsonl | 332aa2a063cf285a584bc1630f9164c15d18ea904b79039c39585cf9919a277d | 24,414 JSONL rows |
This HTML is a derived, static rendering. Percentages are rounded for display; conclusions and intervals come from the unrounded JSON values. Local absolute paths, credentials, transcript text, and app identifiers are intentionally excluded.