Arabic speech recognition · technical evaluation

Cohere Transcribe vs Wit.ai through Tafrigh

Accuracy, speed, configuration sensitivity, and operational evidence across Modern Standard Arabic, regional dialect speech, and a Classical Arabic recitation proxy.

Evaluation updated July 12, 2026 · Self-contained report · Primary scoring: lexical-normalized corpus WER

Executive findings

Evaluation corpus24,414 clips36.393 decoded hours
Overall lexical WER31.32% vs 34.00%Cohere vs Wit/Tafrigh
End-to-end wall28m 25sCohere; Wit 1h 13m 17s
Wit relative time2.578×157.83% more wall time
Headline: Wit/Tafrigh was +2.68 pp worse overall; the paired 95% CI was [+1.67, +4.16] pp. It was significantly better on the aggregate dialect domain and SADA22, statistically tied on Casablanca, but significantly worse on MSA and the Classical Arabic proxy. Cohere completed the same clip suite 2.578× sooner.

A negative delta favors Wit because every delta is Wit minus Cohere. Confidence intervals that exclude zero are labelled decisive for this benchmark. Statistical significance does not by itself establish practical significance or generalization beyond these datasets.

Domain WER at a glance

CohereWit/Tafrigh

July 12 production release

Reliable 1.wav default36.485s114.04× realtime; two external runs
Balanced merge WER27.6424%500 clips across all five datasets
CUDA peak6.06 / 6.10 GiBallocated / reserved on RTX 3060 12 GB
Release scriptd1dda35f9d39…SHA-256, exact model revision pinned
Recommended reliable/fast command: python transcribe_optimized.py --language ar --vad silero --vad-merge --alignment segment 1.wav. Static length-sorted batch 24 remains the RTX 3060 default. Upward adaptation and pinned host transfers remain opt-in because neither improved repeated end-to-end timing.
Repeated production timing for 1.wav (1h 09m 21s)
Configuration External runs Median wall RTFx Generation Batches Words Transcript SHA-256
Static batch 24, pageable (production default) 36.58s / 36.39s 36.485s 114.04 22.370s 8/8 9,852 11e07d260a45355f…
Adaptive batch 24→30 (experimental) 36.82s / 36.40s 36.610s 113.65 22.358s 7/7 9,852 51e635a60f44b98a…
Static batch 24, pinned host memory (experimental) 36.29s / 36.61s 36.450s 114.15 22.129s 8/8 9,852 11e07d260a45355f…

The default median is based on 36.58s and 36.39s process-wall observations. The current pipeline used 182 merged model rows from 378 raw speech spans and retained 4080.224s of the 4160.679s recording. Its output contains 9,852 words and 1,198 subtitle cues. The adaptive transcript has the same word count but a different hash and 12 lexical token edits relative to static; batching is not assumed text-invariant under finite-precision GPU execution.

Fresh same-source 1.wav comparison
Engine Wall RTFx Original → compact cues Words Timing boundary
Cohere production default 36.485s 114.038 1,198 → 284 9,852 Local RTX 3060; two-run wall midpoint
Wit.ai through stock Tafrigh 1.7.8 131.79s 31.571 323 → 237 9,779 Cloud; 8 independent Arabic apps; one observation
Direct one-file result: stock Tafrigh with eight independent apps took 3.612× Cohere's wall time, or 95.305s longer. Open the synchronized playback comparison to listen against both transcripts. Both display columns use Tafrigh's exact 30-word compact rule, but no words were normalized or corrected.

This long lecture has no human transcript, so its 9,852 versus 9,779 output-word counts and audible disagreements are not WER. The comparison separates one-file latency from fleet throughput: 64 truly independent app quotas across eight Tafrigh processes can scale different files concurrently, while this 131.79-second measurement is one file through one eight-app process.

July 12 balanced 500-clip production gate
Configuration Wall RTFx Segments Overall WER MSA WER Dialect WER CA proxy WER S/D/I
Prior rounded-boundary Silero (historical) 48.70s 103.34 729 32.4073% 7.3939% 51.4638% 41.4723% 1,039/696/353
Sample-exact Silero, static 24 40.13s 125.41 729 31.2898% 7.5556% 47.7273% 43.0029% 1,041/692/283
Sample-exact Silero, adaptive 40.15s 125.35 729 31.0570% 7.5556% 47.6888% 41.9825% 1,036/695/270
Sample-exact Silero + merge, static 24 41.50s 121.27 508 27.6424% 6.8687% 46.9569% 28.5714% 852/738/191
How to read the WER improvement: sample-exact merged Silero scored 27.6424%, but this probe contains presegmented utterances. Merge reduced the processor rows from 729 to 508 and changed ASR context. It is evidence for this production configuration, not proof that merging always improves arbitrary continuous audio. The full 24,414-clip suite was not rerun after this patch.
Exact final-profile telemetry
Stage Measured time Context
Decode worker 4.523s ffmpeg
Silero VAD worker 5.918s onnx on CPUExecutionProvider
ASR model load 6.384s Overlapped with audio preparation
ASR wall / generation 23.223s / 22.368s 8 batches; 28,277 generated tokens
Feature preparation wait 0.828s 5.703% aggregate frame padding
Host-to-device copies 0.205s Pageable host tensors in the production default
Productionized review findings
Area Release behavior Why it matters
Batch control Persistent OOM cap learning and frame-balanced retry; upward growth stays opt-in Prevents repeated oversized attempts without changing the measured default
Decoder completeness EOS-aware token-limit detection with affected-row-only retry Live two-token test retried one row at 128 and left no unresolved row
VAD Sample-index timestamps, isolated ONNX/JIT loaders, provider/fallback provenance Actual ONNX and JIT spans matched on a real clip
Approximate timing Words distributed across retained raw speech spans Merged silence is preserved as subtitle gaps instead of synthetic speech time
Alignment FP32 accelerator-side stable log-softmax and in-place OOM reduction Five-file real smoke test: 12 valid words, zero fallback segments
I/O Concrete decode backend, streamed FFmpeg PCM, transactional profile/output publication Avoids a full duplicate PCM buffer and ambiguous auto provenance
Compatibility Transformers 5.13.x runtime bound and recorded validation environment Internal model hooks now fail closed outside the validated series
Telemetry Per-batch frames/tokens/padding/timing plus allocated and reserved VRAM The remaining bottleneck can be measured rather than inferred from rounded logs

The reviewer’s warning about TorchAudio forced_align removal was outdated: the operation was retained during the TorchAudio maintenance transition. Production therefore checks matching Torch/TorchAudio release lines and executes an operation smoke test rather than replacing a working dependency. The release passed 54 transcription-focused tests and 71 benchmark tests, Ruff lint/format, and bytecode compilation. Mypy was unavailable, so no mypy result is claimed.

The portable final/ bundle contains the production implementation as transcribe.py with simplified input.txt/input.srt/input.vtt/input.json output names, bundled vectorized Silero ONNX runtime and applicable licenses, portable runtime requirements, a device-oriented setup guide, version/model provenance, and an installation validator. Model weights remain upstream and download at first use after model-access approval.

Release summary artifact: production_release_20260712.json. Production audit: PRODUCTION_REVIEW_20260712.md. Both are hashed in the artifact inventory below.

Evaluation suite

The frozen suite contains one audio path and reference per clip. Exact waveform samples returned by the configured decoder, rather than source metadata alone, provide the timing denominator.

Frozen evaluation dataset composition
Dataset Role Clips Decoded h Pinned revision Source-card license
Casablanca Eight conversational dialects 6,726 7.703 8951b1b88e28 CC BY-NC-ND 4.0
Common Voice 18 Arabic MSA / read speech 10,471 12.657 1a52eefd8259 CC0
FLEURS ar-EG MSA / read speech 428 1.302 70bb2e84b976 CC BY 4.0
Quran ayah CA proxy Classical Arabic recitation proxy 600 3.982 80cad1ab411c Treat as CC BY-NC-SA 4.0
SADA22 Saudi dialects; broadcast/noise/music 6,189 10.749 094fe2c0fe4b CC BY-NC-SA 4.0

Common Voice, SADA22, and Casablanca references were matched to the exact rows published by the Arabic ASR leaderboard at commit 10cf2c8. The Quran subset is 12 evenly spaced blocks of 50 rows from a three-reciter-held-out split. Its canonical verse references make it a Classical Arabic proxy, not an official Quranic leaderboard score.

Decoded duration was 36.393282 hours versus 36.301534 hours in source metadata, a 330.292-second difference dominated by Common Voice MP3 duration reporting. All 24,414 paths decoded successfully and IDs, references, grouping metadata, and audio paths matched across compared configurations.

Other Arabic evaluation options considered

No single corpus covers MSA, regional dialects, noise, code-switching, and Classical Arabic. These candidates remain useful extensions to the completed suite.

Additional evaluation candidates
Candidate Test scale Coverage Status / tradeoff
Open Universal Arabic ASR Leaderboard About 65.9 h SADA, CV18, MASC, MGB-2, Casablanca Best final public comparison; substantially larger and partly gated
MASC clean/noisy About 25.4 h 20+ dialects, MSA, YouTube, varied acoustics Strong expansion; large download and not run here
MGB-2 About 9.6 h Al Jazeera broadcast, interviews, several dialects Gated and research-only
ArzEn About 2.9 h test Spontaneous Egyptian Arabic-English code-switching Access request required
TunSwitch-CS About 25 min test Tunisian Arabic-French-English code-switching Open but small
Private in-domain holdout Recommended 5-10 h Actual production speakers, channels, and noise Best production-validity check; split by speaker and source

Full-suite accuracy

These are corpus-aggregated lexical-normalized rates. S/D/I means substitutions, deletions, and insertions. The bootstrap resamples 14,319 speaker/utterance clusters overall, with speakers namespaced by dataset where speaker IDs exist.

Overall lexical-normalized comparison
Scope Clips Hours Cohere WER Wit WER Delta Paired 95% CI Cohere CER Wit CER WER S/D/I
Cohere / Wit
Exact disagreements Conclusion
Overall 24,414 36.39 31.32% 34.00% +2.68 pp [+1.67, +4.16] pp 14.24% 20.03% 46,576/17,756/7,629
34,499/38,952/4,669
23,178 / 94.94% Cohere better
Lexical-normalized results by dataset
Scope Clips Hours Cohere WER Wit WER Delta Paired 95% CI Cohere CER Wit CER WER S/D/I
Cohere / Wit
Exact disagreements Conclusion
Casablanca 6,726 7.70 48.76% 48.83% +0.07 pp [-0.48, +0.57] pp 18.73% 26.78% 25,846/7,659/2,777
16,781/17,513/2,040
6,523 / 96.98% No decisive difference
Common Voice 18 Arabic 10,471 12.66 5.54% 12.87% +7.32 pp [+6.71, +8.02] pp 1.53% 5.91% 2,495/275/184
4,420/2,144/293
9,766 / 93.27% Cohere better
FLEURS ar-EG 428 1.30 4.75% 19.73% +14.99 pp [+12.60, +17.52] pp 2.15% 14.68% 232/96/52
534/909/137
428 / 100.00% Cohere better
Quran ayah CA proxy 600 3.98 15.35% 41.70% +26.35 pp [+8.37, +53.13] pp 11.07% 38.46% 252/50/914
329/2,835/140
593 / 98.83% Cohere better
SADA22 6,189 10.75 36.14% 34.89% -1.26 pp [-2.02, -0.50] pp 20.00% 21.90% 17,751/9,676/3,702
12,435/15,551/2,059
5,868 / 94.81% Wit better
Lexical-normalized results by language domain
Scope Clips Hours Cohere WER Wit WER Delta Paired 95% CI Cohere CER Wit CER WER S/D/I
Cohere / Wit
Exact disagreements Conclusion
Classical Arabic 600 3.98 15.35% 41.70% +26.35 pp [+8.37, +53.13] pp 11.07% 38.46% 252/50/914
329/2,835/140
593 / 98.83% Cohere better
Dialect 12,758 17.91 42.60% 41.95% -0.64 pp [-1.12, -0.19] pp 19.76% 24.59% 43,169/17,125/6,426
28,895/32,818/4,000
12,241 / 95.95% Wit better
MSA 11,056 14.50 6.17% 13.96% +7.79 pp [+7.15, +8.47] pp 1.95% 7.23% 3,155/581/289
5,275/3,299/529
10,344 / 93.56% Cohere better
Interpretation: Wit is competitive where the suite is dominated by conversational/dialect speech: it gains 1.26 pp on SADA22 and is effectively tied on Casablanca. Its aggregate dialect gain is 0.64 pp. Cohere holds large advantages on Common Voice, FLEURS, and the Quran proxy. Because dataset composition changes the aggregate, the 2.68 pp overall gap should not be treated as a universal Arabic-language ranking.

Normalization-profile audit

Raw

Exact spelling, punctuation, whitespace-sensitive word boundaries, and diacritics. Useful for output fidelity, but sensitive to stylistic conventions.

Leaderboard repo-exact

Reproduces the public evaluator at commit 10cf2c8, including its malformed punctuation-regex behavior. Included as an audit profile.

Leaderboard intended

Applies the punctuation removal documented by the leaderboard plus Arabic diacritic, hamza, and digit mappings.

Lexical normalized

NFKC, combining-mark and tatweel removal, punctuation-to-space normalization, Arabic/Persian letter folding, and digit mapping. This is the headline profile and the closest disclosed approximation to Cohere's scoring.

Overall results under every scoring profile
Scope Profile Cohere WER Wit WER Delta Cohere CER Wit CER Ref words Cohere WER S/D/I Wit WER S/D/I
Overall Raw 49.09% 48.85% -0.24 pp 21.76% 26.24% 230,422 86,285/18,039/8,797 68,504/39,408/4,650
Overall Leaderboard repo-exact 38.03% 38.04% +0.01 pp 16.59% 21.13% 230,421 60,615/18,128/8,887 43,428/39,487/4,730
Overall Leaderboard intended 31.59% 34.19% +2.60 pp 14.45% 20.17% 229,701 47,211/17,820/7,537 34,914/38,929/4,702
Overall Lexical normalized 31.32% 34.00% +2.68 pp 14.24% 20.03% 229,757 46,576/17,756/7,629 34,499/38,952/4,669
Every profile by dataset
Dataset rates under all four profiles
Scope Profile Cohere WER Wit WER Delta Cohere CER Wit CER Ref words Cohere WER S/D/I Wit WER S/D/I
Casablanca Raw 58.09% 59.33% +1.25 pp 22.58% 29.78% 75,085 32,796/8,141/2,677 24,462/18,033/2,056
Casablanca Leaderboard repo-exact 54.55% 55.50% +0.95 pp 21.24% 28.64% 75,085 30,068/8,178/2,714 21,527/18,061/2,084
Casablanca Leaderboard intended 48.87% 48.95% +0.08 pp 18.97% 26.99% 74,375 25,949/7,676/2,723 16,848/17,495/2,063
Casablanca Lexical normalized 48.76% 48.83% +0.07 pp 18.73% 26.78% 74,416 25,846/7,659/2,777 16,781/17,513/2,040
Common Voice 18 Arabic Raw 38.12% 43.13% +5.02 pp 15.73% 18.59% 53,289 18,976/276/1,060 20,550/2,145/291
Common Voice 18 Arabic Leaderboard repo-exact 17.27% 13.73% -3.54 pp 4.39% 6.22% 53,288 7,867/276/1,061 4,876/2,147/294
Common Voice 18 Arabic Leaderboard intended 5.85% 13.11% +7.26 pp 1.78% 6.10% 53,284 2,650/281/188 4,549/2,143/294
Common Voice 18 Arabic Lexical normalized 5.54% 12.87% +7.32 pp 1.53% 5.91% 53,286 2,495/275/184 4,420/2,144/293
FLEURS ar-EG Raw 19.76% 37.30% +17.54 pp 5.24% 18.07% 8,000 1,418/104/59 1,953/898/133
FLEURS ar-EG Leaderboard repo-exact 15.34% 21.84% +6.50 pp 4.17% 15.11% 8,000 1,064/104/59 690/911/146
FLEURS ar-EG Leaderboard intended 4.90% 19.81% +14.91 pp 2.22% 14.74% 7,994 239/98/55 533/905/146
FLEURS ar-EG Lexical normalized 4.75% 19.73% +14.99 pp 2.15% 14.68% 8,007 232/96/52 534/909/137
Quran ayah CA proxy Raw 104.19% 101.34% -2.85 pp 44.33% 64.65% 7,924 7,308/37/911 5,123/2,801/106
Quran ayah CA proxy Leaderboard repo-exact 19.51% 44.50% +24.99 pp 13.02% 39.11% 7,924 574/49/923 551/2,835/140
Quran ayah CA proxy Leaderboard intended 19.07% 44.18% +25.11 pp 11.92% 39.05% 7,924 550/50/911 526/2,835/140
Quran ayah CA proxy Lexical normalized 15.35% 41.70% +26.35 pp 11.07% 38.46% 7,924 252/50/914 329/2,835/140
SADA22 Raw 45.70% 39.49% -6.21 pp 23.33% 23.02% 86,124 25,787/9,481/4,090 16,416/15,531/2,064
SADA22 Leaderboard repo-exact 40.28% 38.76% -1.52 pp 21.92% 22.84% 86,124 21,042/9,521/4,130 15,784/15,533/2,066
SADA22 Leaderboard intended 36.22% 34.91% -1.31 pp 20.11% 21.92% 86,124 17,823/9,715/3,660 12,458/15,551/2,059
SADA22 Lexical normalized 36.14% 34.89% -1.26 pp 20.00% 21.90% 86,124 17,751/9,676/3,702 12,435/15,551/2,059
Every profile by domain
Domain rates under all four profiles
Scope Profile Cohere WER Wit WER Delta Cohere CER Wit CER Ref words Cohere WER S/D/I Wit WER S/D/I
Classical Arabic Raw 104.19% 101.34% -2.85 pp 44.33% 64.65% 7,924 7,308/37/911 5,123/2,801/106
Classical Arabic Leaderboard repo-exact 19.51% 44.50% +24.99 pp 13.02% 39.11% 7,924 574/49/923 551/2,835/140
Classical Arabic Leaderboard intended 19.07% 44.18% +25.11 pp 11.92% 39.05% 7,924 550/50/911 526/2,835/140
Classical Arabic Lexical normalized 15.35% 41.70% +26.35 pp 11.07% 38.46% 7,924 252/50/914 329/2,835/140
Dialect Raw 52.00% 49.46% -2.54 pp 23.34% 26.64% 157,299 57,671/17,415/6,713 40,468/33,318/4,021
Dialect Leaderboard repo-exact 47.56% 47.25% -0.31 pp 21.97% 26.00% 157,299 50,532/17,491/6,789 36,923/33,348/4,051
Dialect Leaderboard intended 42.69% 42.03% -0.67 pp 19.93% 24.70% 156,589 43,341/17,181/6,330 28,985/32,800/4,023
Dialect Lexical normalized 42.60% 41.95% -0.64 pp 19.76% 24.59% 156,630 43,169/17,125/6,426 28,895/32,818/4,000
MSA Raw 35.38% 40.99% +5.61 pp 14.03% 17.92% 65,199 21,306/587/1,173 22,913/3,289/523
MSA Leaderboard repo-exact 17.29% 15.03% -2.26 pp 4.57% 7.56% 65,198 9,509/588/1,175 5,954/3,304/539
MSA Leaderboard intended 6.45% 14.17% +7.72 pp 2.16% 7.39% 65,188 3,320/589/296 5,403/3,294/539
MSA Lexical normalized 6.17% 13.96% +7.79 pp 1.95% 7.23% 65,203 3,155/581/289 5,275/3,299/529

Paired confidence intervals were computed only for the selected lexical-normalized profile. Deltas in the other profile tables are descriptive corpus-rate differences, not separately bootstrapped estimates.

Dialect and variety detail

Labels are retained from their source datasets. “MSA” is a reference-register label, “More than 1 speaker” and “Unknown” are annotation categories, and tiny categories should not be ranked.

Lexical-normalized result by dialect or annotation label
Scope Clips Hours Cohere WER Wit WER Delta Paired 95% CI Cohere CER Wit CER WER S/D/I
Cohere / Wit
Exact disagreements Conclusion
Algeria 843 0.94 63.04% 59.53% -3.51 pp [-5.49, -1.82] pp 24.31% 31.71% 3,702/530/760
2,487/1,759/468
825 / 97.86% Wit better
Classical Arabic 600 3.98 15.35% 41.70% +26.35 pp [+8.37, +53.13] pp 11.07% 38.46% 252/50/914
329/2,835/140
593 / 98.83% Cohere better
Egypt 825 0.99 31.61% 36.69% +5.07 pp [+3.64, +6.41] pp 11.93% 15.46% 2,120/863/208
2,054/1,283/366
781 / 94.67% Cohere better
Egyptian 96 0.09 35.24% 30.64% -4.61 pp [-10.26, +0.52] pp 16.74% 14.54% 185/71/27
151/65/30
85 / 88.54% No decisive difference
Hijazi 809 1.12 31.99% 32.90% +0.91 pp [-0.71, +2.60] pp 15.96% 20.78% 1,816/802/249
1,180/1,613/156
773 / 95.55% No decisive difference
Jordan 848 0.98 29.69% 31.93% +2.24 pp [+1.01, +3.42] pp 8.82% 11.96% 1,894/569/147
1,705/912/190
809 / 95.40% Cohere better
Khaliji 1,150 1.13 37.70% 42.20% +4.49 pp [+2.17, +6.64] pp 18.56% 28.57% 2,257/829/445
1,435/2,338/179
1,081 / 94.00% Cohere better
MSA 11,056 14.50 6.17% 13.96% +7.79 pp [+7.14, +8.48] pp 1.95% 7.23% 3,155/581/289
5,275/3,299/529
10,344 / 93.56% Cohere better
Mauritania 948 0.94 79.75% 84.94% +5.19 pp [+3.81, +6.56] pp 41.53% 73.05% 5,481/2,060/445
1,408/7,016/82
940 / 99.16% Cohere better
More than 1 speaker اكثر من متحدث 1,320 4.80 39.42% 35.86% -3.56 pp [-4.92, -2.32] pp 23.46% 22.22% 7,769/5,682/1,794
6,227/6,527/1,115
1,303 / 98.71% Wit better
Morocco 1,045 1.01 54.50% 43.01% -11.50 pp [-12.78, -10.23] pp 17.41% 17.77% 4,856/1,382/280
2,900/2,084/159
1,020 / 97.61% Wit better
Najdi 1,703 2.07 31.04% 28.31% -2.73 pp [-4.11, -1.45] pp 16.49% 16.54% 3,288/1,499/553
2,096/2,436/338
1,571 / 92.25% Wit better
Notapplicable 167 0.13 44.12% 51.02% +6.90 pp [+2.91, +10.93] pp 22.36% 34.82% 264/86/59
172/292/9
158 / 94.61% Cohere better
Palestine 667 0.98 37.96% 36.26% -1.69 pp [-2.86, -0.54] pp 12.41% 12.87% 2,358/746/241
2,010/918/268
645 / 96.70% Wit better
Shamali 18 0.02 26.73% 30.69% +3.96 pp [-0.91, +12.50] pp 10.62% 18.28% 47/4/3
35/25/2
16 / 88.89% No decisive difference
UAE 813 0.93 40.38% 45.19% +4.81 pp [+3.29, +6.28] pp 12.92% 24.30% 2,429/799/228
1,829/1,863/176
789 / 97.05% Cohere better
Unknown 762 0.83 44.40% 48.54% +4.14 pp [+0.28, +7.58] pp 25.57% 36.35% 1,680/482/517
800/2,003/126
724 / 95.01% Cohere better
Yemen 737 0.94 50.62% 53.19% +2.58 pp [+1.27, +3.88] pp 19.62% 27.42% 3,006/710/468
2,388/1,678/331
714 / 96.88% Cohere better
Yemeni 7 0.01 65.22% 63.04% -2.17 pp [-18.75, +9.43] pp 59.09% 27.73% 17/11/2
18/6/5
7 / 100.00% No decisive difference
Every scoring profile for every dialect label
Dialect rates under all four scoring profiles
Scope Profile Cohere WER Wit WER Delta Cohere CER Wit CER Ref words Cohere WER S/D/I Wit WER S/D/I
Algeria Raw 66.76% 65.20% -1.55 pp 28.29% 33.64% 8,044 4,023/633/714 2,957/1,838/450
Algeria Leaderboard repo-exact 65.10% 63.80% -1.31 pp 27.46% 33.21% 8,044 3,888/634/715 2,836/1,842/454
Algeria Leaderboard intended 62.66% 59.51% -3.15 pp 24.47% 31.96% 7,914 3,694/534/731 2,488/1,754/468
Algeria Lexical normalized 63.04% 59.53% -3.51 pp 24.31% 31.71% 7,919 3,702/530/760 2,487/1,759/468
Classical Arabic Raw 104.19% 101.34% -2.85 pp 44.33% 64.65% 7,924 7,308/37/911 5,123/2,801/106
Classical Arabic Leaderboard repo-exact 19.51% 44.50% +24.99 pp 13.02% 39.11% 7,924 574/49/923 551/2,835/140
Classical Arabic Leaderboard intended 19.07% 44.18% +25.11 pp 11.92% 39.05% 7,924 550/50/911 526/2,835/140
Classical Arabic Lexical normalized 15.35% 41.70% +26.35 pp 11.07% 38.46% 7,924 252/50/914 329/2,835/140
Egypt Raw 44.93% 51.00% +6.07 pp 16.17% 19.41% 10,116 3,448/874/223 3,494/1,292/373
Egypt Leaderboard repo-exact 39.83% 45.35% +5.53 pp 14.73% 17.86% 10,116 2,928/876/225 2,907/1,300/381
Egypt Leaderboard intended 32.50% 37.08% +4.58 pp 12.33% 15.65% 10,093 2,211/862/207 2,094/1,282/366
Egypt Lexical normalized 31.61% 36.69% +5.07 pp 11.93% 15.46% 10,094 2,120/863/208 2,054/1,283/366
Egyptian Raw 42.09% 35.37% -6.72 pp 18.82% 15.79% 803 242/70/26 187/65/32
Egyptian Leaderboard repo-exact 38.85% 34.99% -3.86 pp 17.82% 15.68% 803 216/70/26 184/65/32
Egyptian Leaderboard intended 36.49% 30.76% -5.73 pp 17.10% 14.59% 803 195/71/27 152/65/30
Egyptian Lexical normalized 35.24% 30.64% -4.61 pp 16.74% 14.54% 803 185/71/27 151/65/30
Hijazi Raw 41.07% 36.64% -4.43 pp 19.02% 21.66% 8,963 2,625/795/261 1,519/1,611/154
Hijazi Leaderboard repo-exact 34.90% 35.83% +0.93 pp 17.34% 21.48% 8,963 2,062/800/266 1,444/1,612/155
Hijazi Leaderboard intended 32.11% 32.91% +0.80 pp 16.03% 20.79% 8,963 1,827/805/246 1,181/1,613/156
Hijazi Lexical normalized 31.99% 32.90% +0.91 pp 15.96% 20.78% 8,963 1,816/802/249 1,180/1,613/156
Jordan Raw 40.16% 45.77% +5.61 pp 11.87% 15.66% 8,804 2,801/577/158 2,919/912/199
Jordan Leaderboard repo-exact 34.23% 37.89% +3.66 pp 10.09% 13.51% 8,804 2,273/580/161 2,217/916/203
Jordan Leaderboard intended 29.75% 31.94% +2.18 pp 8.90% 12.02% 8,792 1,899/570/147 1,706/912/190
Jordan Lexical normalized 29.69% 31.93% +2.24 pp 8.82% 11.96% 8,792 1,894/569/147 1,705/912/190
Khaliji Raw 48.64% 44.49% -4.15 pp 23.03% 29.17% 9,366 3,245/808/503 1,653/2,335/179
Khaliji Leaderboard repo-exact 41.22% 43.92% +2.70 pp 20.97% 29.02% 9,366 2,540/813/508 1,598/2,336/180
Khaliji Leaderboard intended 37.84% 42.28% +4.44 pp 18.72% 28.60% 9,366 2,271/830/443 1,443/2,338/179
Khaliji Lexical normalized 37.70% 42.20% +4.49 pp 18.56% 28.57% 9,366 2,257/829/445 1,435/2,338/179
MSA Raw 35.38% 40.99% +5.61 pp 14.03% 17.92% 65,199 21,306/587/1,173 22,913/3,289/523
MSA Leaderboard repo-exact 17.29% 15.03% -2.26 pp 4.57% 7.56% 65,198 9,509/588/1,175 5,954/3,304/539
MSA Leaderboard intended 6.45% 14.17% +7.72 pp 2.16% 7.39% 65,188 3,320/589/296 5,403/3,294/539
MSA Lexical normalized 6.17% 13.96% +7.79 pp 1.95% 7.23% 65,203 3,155/581/289 5,275/3,299/529
Mauritania Raw 82.29% 86.61% +4.31 pp 44.85% 73.70% 10,110 5,774/2,130/416 1,566/7,108/82
Mauritania Leaderboard repo-exact 81.22% 85.87% +4.65 pp 43.59% 73.43% 10,110 5,651/2,137/423 1,491/7,108/82
Mauritania Leaderboard intended 79.66% 84.97% +5.31 pp 41.67% 73.11% 10,004 5,475/2,071/423 1,410/7,007/83
Mauritania Lexical normalized 79.75% 84.94% +5.19 pp 41.53% 73.05% 10,014 5,481/2,060/445 1,408/7,016/82
More than 1 speaker اكثر من متحدث Raw 48.53% 42.46% -6.07 pp 26.80% 23.83% 38,672 11,204/5,539/2,025 8,782/6,518/1,121
More than 1 speaker اكثر من متحدث Leaderboard repo-exact 44.79% 41.74% -3.05 pp 25.81% 23.66% 38,672 9,721/5,557/2,043 8,504/6,518/1,121
More than 1 speaker اكثر من متحدث Leaderboard intended 39.50% 35.88% -3.62 pp 23.58% 22.24% 38,672 7,786/5,698/1,790 6,234/6,527/1,115
More than 1 speaker اكثر من متحدث Lexical normalized 39.42% 35.86% -3.56 pp 23.46% 22.22% 38,672 7,769/5,682/1,794 6,227/6,527/1,115
Morocco Raw 63.23% 53.95% -9.28 pp 20.78% 20.73% 11,959 5,921/1,374/267 4,184/2,076/192
Morocco Leaderboard repo-exact 62.20% 52.17% -10.03 pp 20.08% 20.05% 11,959 5,794/1,376/269 3,971/2,076/192
Morocco Leaderboard intended 54.43% 43.01% -11.41 pp 17.42% 17.83% 11,959 4,855/1,386/268 2,901/2,084/159
Morocco Lexical normalized 54.50% 43.01% -11.50 pp 17.41% 17.77% 11,959 4,856/1,382/280 2,900/2,084/159
Najdi Raw 40.64% 31.57% -9.07 pp 19.37% 17.29% 17,201 4,921/1,483/587 2,662/2,432/337
Najdi Leaderboard repo-exact 33.99% 30.59% -3.40 pp 17.71% 17.05% 17,201 3,758/1,492/596 2,493/2,432/337
Najdi Leaderboard intended 31.02% 28.34% -2.69 pp 16.55% 16.55% 17,201 3,292/1,512/532 2,100/2,436/338
Najdi Lexical normalized 31.04% 28.31% -2.73 pp 16.49% 16.54% 17,201 3,288/1,499/553 2,096/2,436/338
Notapplicable Raw 49.84% 53.29% +3.45 pp 24.05% 35.46% 927 316/84/62 193/292/9
Notapplicable Leaderboard repo-exact 45.85% 52.97% +7.12 pp 22.90% 35.40% 927 279/84/62 190/292/9
Notapplicable Leaderboard intended 44.55% 51.13% +6.58 pp 22.47% 34.86% 927 268/86/59 173/292/9
Notapplicable Lexical normalized 44.12% 51.02% +6.90 pp 22.36% 34.82% 927 264/86/59 172/292/9
Palestine Raw 47.99% 51.12% +3.12 pp 15.36% 16.95% 8,833 3,254/752/233 3,314/920/281
Palestine Leaderboard repo-exact 43.51% 44.62% +1.11 pp 13.97% 15.06% 8,833 2,844/759/240 2,732/924/285
Palestine Leaderboard intended 37.96% 36.30% -1.66 pp 12.50% 13.00% 8,812 2,358/746/241 2,014/917/268
Palestine Lexical normalized 37.96% 36.26% -1.69 pp 12.41% 12.87% 8,813 2,358/746/241 2,010/918/268
Shamali Raw 36.14% 33.66% -2.48 pp 13.21% 18.80% 202 63/4/6 41/25/2
Shamali Leaderboard repo-exact 30.69% 31.19% +0.50 pp 11.85% 18.49% 202 52/4/6 36/25/2
Shamali Leaderboard intended 26.73% 30.69% +3.96 pp 10.73% 18.28% 202 47/4/3 35/25/2
Shamali Lexical normalized 26.73% 30.69% +3.96 pp 10.62% 18.28% 202 47/4/3 35/25/2
UAE Raw 52.70% 59.02% +6.32 pp 17.52% 28.54% 8,766 3,428/969/223 2,969/2,045/160
UAE Leaderboard repo-exact 49.41% 53.35% +3.95 pp 16.29% 26.84% 8,766 3,131/973/227 2,460/2,051/166
UAE Leaderboard intended 40.38% 45.24% +4.86 pp 13.20% 24.54% 8,559 2,435/799/222 1,833/1,863/176
UAE Lexical normalized 40.38% 45.19% +4.81 pp 12.92% 24.30% 8,559 2,429/799/228 1,829/1,863/176
Unknown Raw 54.41% 50.98% -3.43 pp 29.76% 37.01% 6,034 2,238/481/564 949/2,001/126
Unknown Leaderboard repo-exact 47.51% 50.61% +3.10 pp 27.81% 36.92% 6,034 1,818/483/566 927/2,001/126
Unknown Leaderboard intended 44.45% 48.56% +4.11 pp 25.78% 36.38% 6,034 1,689/488/505 801/2,003/126
Unknown Lexical normalized 44.40% 48.54% +4.14 pp 25.57% 36.35% 6,034 1,680/482/517 800/2,003/126
Yemen Raw 64.14% 61.75% -2.39 pp 25.53% 30.67% 8,453 4,147/832/443 3,059/1,842/319
Yemen Leaderboard repo-exact 57.45% 60.07% +2.63 pp 23.13% 30.07% 8,453 3,559/843/454 2,913/1,844/321
Yemen Leaderboard intended 51.13% 53.76% +2.63 pp 20.43% 28.19% 8,242 3,022/708/484 2,402/1,676/353
Yemen Lexical normalized 50.62% 53.19% +2.58 pp 19.62% 27.42% 8,266 3,006/710/468 2,388/1,678/331
Yemeni Raw 71.74% 67.39% -4.35 pp 62.73% 28.64% 46 21/10/2 20/6/5
Yemeni Leaderboard repo-exact 65.22% 67.39% +2.17 pp 60.00% 28.64% 46 18/10/2 20/6/5
Yemeni Leaderboard intended 65.22% 63.04% -2.17 pp 59.55% 27.73% 46 17/11/2 18/6/5
Yemeni Lexical normalized 65.22% 63.04% -2.17 pp 59.09% 27.73% 46 17/11/2 18/6/5

Timing and throughput

Primary equivalent timing boundary
Engine/config Wall Decoded RTFx Clips/s Scope
Optimized Cohere BF16 length batch 24 28m 25s 76.82 14.31 Local GPU, all-samples wall
Wit.ai via Tafrigh default 1h 13m 17s 29.79 5.55 Cloud API + Tafrigh preprocessing, cumulative active wall
Observed speed difference: Wit took 2.578× as long, adding 44m 52s. Its throughput was 0.388× Cohere's, or 61.21% lower. Each wall time is one observation; no timing confidence interval was estimated.
Timing component provenance
Component Time Interpretation
Cohere audio decode 16m 08s Serialized aggregate component
Cohere feature processing 0m 50s Serialized aggregate component
Cohere model generation 11m 21s Generation-only RTFx 192.25; not equivalent to Wit wall
Cohere text decode 0.425s Aggregate
Wit API request time sum 20h 42m 33s Sum across concurrent requests; exceeds wall by design
Wit preprocessing worker time sum 1h 53m 09s Sum across concurrent workers; not wall
Wit active wall 1h 13m 17s Recovered invocations combined

Cohere generation alone was 11m 21s, but excluding decode and feature preparation makes it an invalid direct boundary against Wit end-to-end wall. Wit active time excludes credential entry, validation, debugging downtime, exact-duration analysis, and WER scoring. Its first recovered component is accurate to roughly one second; later components came from sanitized ledgers.

Continuous 1.wav case study

Observed processing of 1h 09m 21s input audio
Run Evidence Wall RTFx Key stages Output
July 12 production: Silero + merge + segment Two /usr/bin/time runs + exact profiles 36.485s median 114.04 22.37s generation; 8 batches; 6.06/6.10 GiB allocated/reserved 9,852 words; 1,198 cues
Original transcribe.py Exact stdout 2m 38s 26.3 62s ASR; 44s emissions/alignment 9,754 words; 1,225 cues
July 11 historical: plain transcript only Exact stdout + /usr/bin/time 50.61s external 82.2 24s ASR; alignment/cues skipped; 6.06 GiB allocated peak 9,755 words; 378 text lines
July 11 historical: segment interpolation Exact stdout + /usr/bin/time 49.23s external 84.5 23s ASR; alignment skipped; 6.06 GiB allocated peak 9,755 words; 1,131 cues
July 11 historical: FP16 word alignment Exact stdout + /usr/bin/time 68.41s external 60.8 23s ASR; 16s emissions; 6.06 GiB allocated peak 9,755 words; 1,225 cues
July 11 historical: FP32 word alignment Exact stdout + /usr/bin/time 97.56s external 42.7 23s ASR; 46s emissions; 6.06 GiB allocated peak 9,755 words; 1,225 cues
Native F16, E4/D8 argmax Guarded release harness 84.36s external 49.3 36.9s encoder; 40.6s decoder; 6,900 MiB peak 9,881 words; text only
Native Q8_0, E4/D8 argmax Guarded release harness 60.55s external 68.7 31.3s encoder; 22.6s decoder; 5,144 MiB peak 9,881 words; text only
Native Q4_K imatrix, E4/D8 argmax Guarded release harness 53.17s external 78.3 31.8s encoder; 14.9s decoder; 4,264 MiB peak 9,882 words; text only

In the frozen July 11 timestamp experiment, all four output/timestamp modes shared the same ASR inputs and produced the same joined text. That statement does not compare later sample-exact or merged segmentation policies. --text-only writes one non-empty VAD/ASR segment per line and skips word allocation, cue grouping, and MMS alignment; its historical 50.61-second wall was effectively tied with segment mode because ASR and VAD dominated. A frozen, word-indexed rerun paired 9,755 words between FP16 and FP32. The combined-boundary absolute median/p95/p99 were 0.00/0.00/0.00 ms, and the maximum was 1420.00 ms. This supersedes the earlier ordinal-SRT comparison: cue reflow changed which words an ordinal cue contained, so its apparent 3.12-second maximum was not a valid acoustic-boundary comparison. FP32 remains the implementation-reference word-alignment mode; --alignment segment is the production reliable/fast recommendation when exact word boundaries are unnecessary.

The native rows use the reference quiet-boundary planner rather than Silero and write no timestamps. Native Q8 differed from native F16 by 28 lexical word edits (0.2834%); Q4 differed by 116 (1.1740%). Because 1.wav has no human transcript, these are model disagreements, not WER. The frozen 500-clip reference benchmark, not this file, determines the precision recommendation.

VAD and text-only segmentation modes

--text-only removes timestamp and subtitle work; --vad controls which waveform regions reach the ASR model. Changing VAD can therefore change recognition text and accuracy, even though every mode writes the same plain-text output format. The production default remains 30-second maximum spans. The pinned processor accepts one row through 35 seconds, and the script verifies the actual processor row count at runtime; 35 seconds is supported as an experiment, not presumed faster or more accurate.

Continuous 1.wav text-only segmentation comparison
Mode Options External wall RTFx Segments Audio retained Words Normalized distance vs Silero
Silero ONNX --text-only --vad silero 50.61s 82.21 378 95.75% 9,755 Reference mode
Auditok, threshold 50 --text-only --vad auditok --energy-threshold 50 36.45s 114.15 144 99.95% 9,863 9.0909% (886 edits)
No VAD, fixed 30 s --text-only --vad none --max-dur 30 34.97s 118.98 139 100.00% 9,905 9.1627% (893 edits)
No long-form accuracy reference: 1.wav has no human reference transcript. Auditok and no-VAD were 9.09% and 9.16% different from Silero by normalized token edit distance. This roughly 9% result is disagreement, not accuracy, and does not identify which transcript is correct.
Balanced 500-clip end-to-end text-only runs
Mode External wall RTFx Run VAD worker Standalone VAD audit Segments Audio retained No segments Lexical WER WER delta vs Silero S/D/I
Silero ONNX 48.70s 103.40 27s 21.49s 729 69.25% 14 32.4073% Baseline 1,039/696/353
Auditok, threshold 50 49.75s 101.22 4s 1.71s 723 88.79% 1 28.4184% -3.9888 pp 948/537/346
No VAD, fixed 30 s 49.58s 101.57 0s 0.00s 542 100.00% 0 23.2656% -9.1417 pp 835/384/280

The same already-segmented corpus scored 22.9551% WER on the direct ASR benchmark path, outside this end-to-end script timing boundary. Silero completed in 48.70s, Auditok in 49.75s, and no-VAD in 49.58s. The cheaper VAD stages did not reduce complete folder wall in this single observation because preprocessing overlaps model loading/GPU work and each segmentation policy changed the ASR inputs.

VAD-mode lexical WER by dataset
Dataset/domain Silero WER Auditok 50 WER No-VAD 30 s WER Auditok delta No-VAD delta
Casablanca 55.8099% 52.2007% 50.7923% -3.6092 pp -5.0176 pp
Common Voice 18 Arabic 5.4688% 8.2031% 5.4688% +2.7344 pp +0.0000 pp
FLEURS ar-EG 7.8961% 17.2695% 6.6225% +9.3734 pp -1.2736 pp
Quran ayah CA proxy 41.4723% 19.0962% 15.1603% -22.3761 pp -26.3120 pp
SADA22 48.0822% 40.7534% 38.0822% -7.3288 pp -10.0000 pp
VAD-mode lexical WER by language domain
Dataset/domain Silero WER Auditok 50 WER No-VAD 30 s WER Auditok delta No-VAD delta
Classical Arabic 41.4723% 19.0962% 15.1603% -22.3761 pp -26.3120 pp
Dialect 51.4638% 45.7627% 43.6441% -5.7011 pp -7.8197 pp
MSA 7.3939% 15.3939% 6.3838% +8.0000 pp -1.0101 pp
Boundary validity: the 500-file probe consists of presegmented evaluation utterances. Its lower no-VAD WER shows that destructive trimming can be avoided on clips that already have useful boundaries; it cannot prove continuous hard-boundary accuracy. On continuous recordings, fixed 30-second windows can split words and may transcribe or hallucinate silence. Auditok threshold 76 was recording-specific and rejected: on the diverse probe it retained only 18.85% of audio and returned no segments for 234 files.

Silero remains the accuracy-conservative default for arbitrary continuous audio. Auditok threshold 50 and --vad none --max-dur 30 are explicit speed experiments whose suitability should be measured on representative, continuously recorded ground truth.

Timestamp-mode benchmark

Balanced production probe500 clips1.3988 decoded hours
Transcript parity100.00%6,119 words in every mode
FP16 within 20 ms99.665%12,238 word boundaries vs FP32
Segment median drift260 msFP32 implementation reference
Reference boundary: these values measure disagreement with FP32 MMS forced alignment, not absolute acoustic error. Common Voice, FLEURS, SADA22, Casablanca, and the Quran sample provide clip boundaries or durations, but no human word-level timestamps. A defensible absolute-error claim requires a manually annotated word-timing set.
Complete 500-file production runs
Mode External wall RTFx Wall vs segment CTC emissions Words Cues Fallback segments Lexical WER
Plain transcript only 48.70s 103.40 1.00× Skipped 6,119 N/A 0 32.4073%
Segment interpolation 48.52s 103.79 1.00× Skipped 6,119 1,032 0 32.4073%
FP16 forced word alignment 119.07s 42.29 2.45× 66.00s 6,119 1,194 2 32.4073%
FP32 forced word alignment 233.26s 21.59 4.81× 177.00s 6,119 1,192 2 32.4073%

--text-only completed in 48.70s versus 48.52s for segment timing: a 0.18-second difference inside run-to-run variation. It produces only .aligned.txt, with one non-empty VAD/ASR segment per line. The periodic collection fix discovered during this benchmark reduced FP16 wall from 192.24s to 119.07s (38.1%) and FP32 from 302.15s to 233.26s (22.8%). Full heap collection now runs every 64 aligned files instead of after every file.

Frozen-input timing stage only (one shared decode cache)
Mode Shared decode Model load Timing compute Cold equivalent Compute RTFx Alignment peak GiB
Segment interpolation 20.23s 0.00s 0.01s 20.26s 399167.81 N/A
FP16 forced word alignment 20.23s 2.03s 66.64s 88.92s 75.56 2.26
FP32 forced word alignment 20.23s 1.60s 178.00s 199.85s 28.29 4.10

On identical frozen audio, boundaries, and ASR text, FP16 timing compute was 2.67× faster than FP32. The complete production ratio is smaller because both modes pay the same ASR, VAD, startup, decode, and output costs.

Word-indexed drift against the FP32 implementation reference
Comparison Abs mean ms Median p95 p99 Max ≤20 ms ≤100 ms >500 ms >1 s Exact word intervals Matched cue spans
Segment interpolation vs FP32 414.05 260.00 1331.85 2289.20 7760.00 12.706% 27.341% 28.1255% 9.5604% 1.59% 610/1,192
FP16 CTC vs FP32 0.91 0.00 0.00 0.00 1740.00 99.665% 99.853% 0.0327% 0.0163% 99.10% 1,190/1,192

FP16 changed only a small tail: 99.665% of word boundaries were within 20 ms and 99.853% were within 100 ms, but four boundaries exceeded 500 ms and two exceeded one second. Both one-second outliers were in the Quran proxy, where a repeated phrase selected a different CTC path. Median, p95, and p99 alone would hide that case. Segment timing is intentionally approximate: it spreads words uniformly inside each VAD segment.

Timestamp drift by dataset (milliseconds)
Dataset Segment median Segment p95 Segment max FP16 median FP16 p95 FP16 p99 FP16 max FP16 >500 ms
Common Voice 18 Arabic 180.00 560.00 1410.00 0.00 0.00 0.00 60.00 0
FLEURS ar-EG 266.67 970.96 7760.00 0.00 0.00 0.00 60.00 0
SADA22 240.00 1403.86 6370.00 0.00 0.00 0.00 80.00 0
Casablanca 172.62 796.81 2640.00 0.00 0.00 0.00 80.00 0
Quran ayah CA proxy 580.00 1860.01 3546.67 0.00 0.00 60.00 1740.00 4
Timestamp drift by language domain (milliseconds)
Domain Segment median Segment p95 Segment max FP16 median FP16 p95 FP16 p99 FP16 max FP16 >500 ms
MSA 244.88 920.00 7760.00 0.00 0.00 0.00 60.00 0
Dialect 206.67 1126.80 6370.00 0.00 0.00 0.00 80.00 0
Classical Arabic 580.00 1860.01 3546.67 0.00 0.00 60.00 1740.00 4
Recognition accuracy is unchanged by downstream timing mode
Dataset/domain Text-only WER Segment WER FP16 WER FP32 WER
Common Voice 18 Arabic 5.4688% 5.4688% 5.4688% 5.4688%
FLEURS ar-EG 7.8961% 7.8961% 7.8961% 7.8961%
SADA22 48.0822% 48.0822% 48.0822% 48.0822%
Casablanca 55.8099% 55.8099% 55.8099% 55.8099%
Quran ayah CA proxy 41.4723% 41.4723% 41.4723% 41.4723%
MSA 7.3939% 7.3939% 7.3939% 7.3939%
Dialect 51.4638% 51.4638% 51.4638% 51.4638%
Classical Arabic 41.4723% 41.4723% 41.4723% 41.4723%

All 500 hypotheses and normalized texts matched across all four modes; the three timing modes also matched every structured word key. FP16 and FP32 each used uniform fallback timing for the same 97 words in two pathological repetition segments; fallback rows are included in the overall comparison and separately identified in the machine-readable report. Cue timing was matched by underlying word span, never by cue ordinal.

Cohere optimization evidence

Four complete BF16 batch/order configurations were evaluated on the same 24,414 clips. Length-sorted batch 24 was selected as the foundation. The later projection cache and conservative repetition-loop stop produced the final Cohere result used throughout the engine comparison.

Full-suite Cohere batch configuration comparison
Config Order Batch Wall Wall RTFx Generation Gen RTFx WER CER Δ vs ordered Paired 95% CI Peak GiB OOM splits
bf16_ordered_b24 ordered 24 48m 54s 44.66 28m 44s 76.01 32.301% 16.429% Baseline Not estimated 9.75 0
bf16_length_b24 length 24 33m 26s 65.32 14m 51s 147.00 32.258% 16.386% -0.043 pp [-0.21, +0.08] pp 10.18 1
bf16_length_b32 length 32 33m 48s 64.60 15m 12s 143.69 32.413% 16.546% +0.112 pp [-0.14, +0.35] pp 11.03 1
bf16_length_b16 length 16 34m 11s 63.87 15m 42s 139.07 32.284% 16.423% -0.017 pp [-0.29, +0.22] pp 9.85 0
Final projection-cache and repetition-stop result
Config Wall Wall RTFx Generation Gen RTFx WER CER WER delta Paired 95% CI
bf16_length_b24 33m 26s 65.32 14m 51s 147.00 32.258% 16.386% Baseline Not estimated
bf16_length_b24_projcache_repstop 28m 25s 76.82 11m 21s 192.25 31.320% 14.241% -0.937 pp [-1.48, -0.48] pp
Long-form ASR-stage approach probe
Approach Segments Total Prepare Generate Peak alloc GB Peak reserved GB
bf16_chron_full 365 64.64s 12.16s 52.20s 6.52 11.44
bf16_sorted_full 365 26.61s 3.98s 22.55s 6.47 6.95
bf16_rep_cpu_features 96 8.82s 1.50s 7.30s 6.52 6.95
bf16_rep_gpu_features 96 8.65s 1.33s 7.28s 6.52 6.95
bf16_rep_eager 96 16.08s 1.30s 14.73s 6.52 6.95
fp16_sorted_full 365 34.26s 3.95s 30.24s 6.48 12.08

Production-pipeline changes

Implemented production optimizations
Area Measured or functional result Implementation boundary
Silero VAD runtime 5.918s worker time on 1.wav Vectorized ONNX CPU inference, sample-index boundaries, isolated TorchScript fallback, and provider provenance
Feature pipeline 26.0s to 22.2s in the sorted ASR pass Prepare one batch ahead while the GPU generates; decoded text was identical in this probe
Corpus batching One ASR and one aligner load per input set Global stable segment identities and optional duration sorting across files
Alignment memory Roughly 800 MB less redundant allocation on an hour-scale file Construct CTC windows a batch at a time; input windows and ordered outputs were parity-checked
Resource controls Bounded decoded-audio groups and worker queues Serialized GPU work, conservative CPU preprocessing, persistent OOM cap learning, and frame-balanced split/retry
Decoder completeness Missing EOS at the token ceiling is no longer silent Retry only affected rows up to the model positional cap and record unresolved rows prominently
Decoder hot path Projection/mask work cached across autoregressive steps Model-specific hooks are guarded by the validated Transformers 5.13.x compatibility bound
Batch-file behavior Multiple paths and recursive folders Collision checks, preserved relative paths, per-file failure isolation, and one output set per source
Alignment cleanup FP16 folder wall 192.24s to 119.07s Run full Python GC every 64 files instead of after every short clip
Alignment numerics Finite FP32 log probabilities under extreme-logit regression Crop on accelerator and use torch.log_softmax(logits.float()) before host transfer
Viterbi safety 500/500 files completed after a native signal-11 failure Use TorchAudio's maintained CTC forced-align operation instead of the package's unsafe ctypes buffer wrapper
Merged subtitle timing Known pauses remain visible Retain raw VAD speech spans and distribute approximate words over speech rather than intervening silence
Decode memory/provenance FFmpeg streams into one backing buffer Record the concrete TorchCodec, librosa, or FFmpeg backend and reuse it for alignment
Output publication Transactional TXT/SRT/VTT/JSON writes Raw indexed words, cues, rollback, temporary cleanup, permission preservation, and source-mutation detection
Optimization telemetry 6.06/6.10 GiB allocated/reserved Exact per-batch rows, frames, padding, tokens, transfers, generation, OOM, and truncation events in atomic JSON

Rejected or non-default trials: torch.compile recompiled dynamic shapes and hurt this one-shot workload; FP16 ASR was slower and changed text; eager attention, oversized adaptive batches, pinned host memory, and shorter CTC variants did not provide a reliable measured win. FP16 alignment remains available explicitly, but it can shift word timestamps and is therefore not the parity default. The actual decoder is selected per installed environment and recorded in profile provenance; the final 1.wav gate resolved to FFmpeg.

Native GGUF batch engine

The new cohere-batch executable is a separate C++/GGML research runtime. It loads one GGUF model for every supplied file and directory, plans long-audio windows globally, computes deterministic CPU features in parallel, batches encoder and ragged decoder work, can select greedy tokens on the GPU, and writes collision-safe plain-text outputs atomically. It does not use Silero, forced alignment, or the Python model implementation.

Frozen 500-clip native precision gate
Model Internal wall External wall Steady RTFx Load Encoder Decoder Peak VRAM Lexical WER Delta vs F16 Paired 95% CI Lexical changes
Native F16 109.866s 110.204s 46.19 0.853s 46.265s 57.528s 7,278 MiB 22.629% Reference Reference 0 / 500
Native Q8_0 76.920s 77.330s 65.94 0.546s 39.834s 31.202s 5,062 MiB 22.784% +0.155 pp [-0.045, +0.437] pp 20 / 500
Native Q4_K imatrix 63.160s 63.414s 80.21 0.382s 40.099s 17.345s 4,188 MiB 23.157% +0.528 pp [-0.315, +1.609] pp 153 / 500
Native lexical WER and paired delta by language domain
Domain F16 Q8_0 Q4_K imatrix
MSA 6.465% 6.424% (-0.040 pp; CI [-0.131, +0.000]) 6.626% (+0.162 pp; CI [-0.165, +0.511])
Dialect 42.604% 43.028% (+0.424 pp; CI [-0.043, +1.120]) 44.299% (+1.695 pp; CI [-0.112, +4.199])
Classical Arabic 13.994% 13.994% (+0.000 pp; CI [+0.000, +0.000]) 12.974% (-1.020 pp; CI [-3.066, +0.134])
Native lexical WER by dataset
Dataset F16 Q8_0 Q4_K imatrix
Common Voice 18 Arabic 5.859% 5.664% (-0.195 pp) 6.445% (+0.586 pp)
FLEURS ar-EG 6.623% 6.623% (+0.000 pp) 6.673% (+0.051 pp)
SADA22 36.849% 37.329% (+0.479 pp) 37.671% (+0.822 pp)
Casablanca 50.000% 50.352% (+0.352 pp) 52.817% (+2.817 pp)
Quran ayah CA proxy 13.994% 13.994% (+0.000 pp) 12.974% (-1.020 pp)
Same frozen 500 source clips under each text-only production boundary
Engine/config External wall RTFx Peak GPU Lexical WER Boundary
Python BF16, fixed 30 s, no VAD 49.58s 101.57 6.07 GiB 23.266% HF runtime; plain text; fixed 30 s planner
Native F16 110.20s 45.67 7.11 GiB 22.629% GGUF runtime; plain text; reference quiet-boundary planner
Native Q8_0 77.33s 65.08 4.94 GiB 22.784% GGUF runtime; plain text; reference quiet-boundary planner
Native Q4_K imatrix 63.41s 79.36 4.09 GiB 23.157% GGUF runtime; plain text; reference quiet-boundary planner
Wit.ai through Tafrigh, 8 apps 123.50s 40.78 Cloud 34.472% Auditok, MP3 payloads, cloud requests
Decision: Q8 is the balanced native default. It was 1.428x faster externally than native F16, saved 30.4% peak VRAM, and changed 20/500 lexical hypotheses. Its +0.155 pp WER point estimate had a paired utterance-bootstrap 95% interval of [-0.045, +0.437] pp. Q4 is a throughput tier, not the reliability default: it changed 153/500 lexical hypotheses and its dialect WER increased by 1.695 pp.

The cross-engine table uses the same 500 source clips and references, but it is not a controlled decoder-only comparison: the Python row uses fixed 30-second no-VAD chunks, native uses the pinned quiet-boundary planner up to 35 seconds, and Tafrigh uses Auditok plus MP3 cloud requests. On this workload the established Python BF16 path remained faster than every native variant. Native Q8 is therefore an implementation and deployment alternative, not a new overall speed record.

Upstream contribution opportunities

The final script was mapped against current default-branch source in four upstream repositories. GitHub searches covered open and closed issues and pull requests, exact symbols, broader concepts, and nearby changes. “No duplicate found” means no matching public indexed work was found as of July 12, 2026; it cannot exclude an unlinked draft or private branch.

Worthwhile upstream work, ordered by expected value
Rank Candidate Target Evidence Duplicate check and action
1 · P0 Generic repeated-block stopping criterion Transformers generation Local per-row sticky criterion stopped pathological loops; 500-clip wall 59.603→51.067s after projection cache and WER improved on the full suite Existing issue #32902; no linked/searchable PR. Comment with ASR evidence and contribute there.
2 · P0 Project Cohere encoder states once per autoregressive decode Transformers Cohere ASR Current decoder repeats an invariant projection per generated token; local ablation 61.208→59.603s wall, 0/500 text changes Issue search / PR search: no open/closed duplicate. #45214 only handled device placement.
3 · P0 Stable ONNX emission log-softmax deskpai/ctc_forced_aligner Current NumPy log(exp(x)/sum(exp(x))) overflows; extreme-logit regression is non-finite while the shifted/log-softmax form is finite Issue search / PR search: no duplicate; submit a focused issue and PR.
4 · P1 Strict Cohere chunk-text reassembly Transformers Cohere processor Reject mismatched/duplicate/gapped metadata and handle all-empty chunks instead of silent zip truncation or IndexError Issue search / PR search: no duplicate; suitable small processor PR.
5 · P1 Prepare Cohere cross-attention mask once Transformers Cohere ASR Current decoder rebuilds an invariant mask per token; local 500-clip generation 28.533→28.171s, 0/500 text changes Merged #46738 is adjacent infrastructure, not a duplicate; current main still repeats the work.
6 · P1 State-preserving vectorized offline Silero ONNX snakers4/silero-vad Batch sequence frames while carrying recurrent state; exact ONNX/JIT span parity locally and ~5.8s VAD worker time for 4,161s audio Concept exists in discussion #408 and #217, but no implementation PR. Coordinate before coding.
7 · P1 Bounded streaming CTC emissions and learned OOM cap MahmoudAshraf97 + DeskPai aligners Normalize/crop each batch, copy directly to bounded host storage, preserve completed windows after OOM, and reuse the learned cap Open issue #89 reports long-file failures; no streaming PR. Reply with a profile/reproducer first.
Existing urgent work, do not duplicate: ctc-forced-aligner issue #94 already documents native backpointer heap corruption; contribute the bounds/allocation fix there. Closed PR #90 overlaps the emission-derived <star> vocabulary index; only revive it with a new current-main failure fixture.

Recommended contribution order: first the stable DeskPai log-softmax because it is a tiny correctness patch; then the Cohere projection change with a profiler/call-count test; then the generic repetition criterion on the already accepted feature request. Keep the projection and mask patches separate so training, direct decoder calls, beam expansion, compilation, and multi-device behavior can be reviewed independently. For Silero, contribute reproducible export code and exact 8/16 kHz parity tests rather than only the ONNX binary.

The complete audit, pinned upstream commit hashes, exact local line mappings, search queries, neighboring PR analysis, and proposed tests are in UPSTREAM_CANDIDATES_20260712.md.

Wit/Tafrigh configuration probe

The duration-stratified probe contains 500 clips, 100 from each dataset (1.398810 decoded hours). It tests Tafrigh-compatible 15/16-second cuts, its actual generated-noise padding, true silence, and no padding. The baseline is exact Tafrigh v1.7.8 behavior.

Overall 500-clip Wit configuration probe
Configuration Wall RTFx Requests WER CER Δ vs default Paired 95% CI Exact disagreements
15 s, Tafrigh noise (default) 123.50s 40.78 894 34.47% 24.91% Baseline Not estimated 0 / 0.00%
15 s, true silence 125.52s 40.12 894 34.07% 23.89% -0.40 pp [-2.25, +1.65] pp 241 / 48.20%
16 s, Tafrigh noise 123.01s 40.94 880 33.93% 24.57% -0.54 pp [-1.23, +0.07] pp 151 / 30.20%
16 s, true silence 124.38s 40.49 880 33.88% 23.93% -0.59 pp [-2.44, +1.65] pp 232 / 46.40%
16 s, no padding 121.57s 41.42 880 42.70% 33.22% +8.23 pp [+1.13, +18.40] pp 354 / 70.80%
Probe configurations under all scoring profiles
Configuration Profile WER CER WER Δ vs default WER S/D/I
15 s, Tafrigh noise (default) Raw 57.30% 36.98% Baseline 2,198/1,398/98
15 s, Tafrigh noise (default) Leaderboard repo-exact 37.47% 25.65% Baseline 902/1,407/107
15 s, Tafrigh noise (default) Leaderboard intended 35.01% 25.08% Baseline 745/1,402/107
15 s, Tafrigh noise (default) Lexical normalized 34.47% 24.91% Baseline 714/1,403/104
15 s, true silence Raw 55.64% 36.07% -1.66 pp 2,174/1,308/105
15 s, true silence Leaderboard repo-exact 37.02% 24.59% -0.45 pp 952/1,319/116
15 s, true silence Leaderboard intended 34.57% 24.05% -0.43 pp 797/1,312/117
15 s, true silence Lexical normalized 34.07% 23.89% -0.40 pp 768/1,313/114
16 s, no padding Raw 57.61% 43.84% +0.31 pp 1,667/1,964/83
16 s, no padding Leaderboard repo-exact 45.34% 33.85% +7.86 pp 864/1,970/89
16 s, no padding Leaderboard intended 42.94% 33.37% +7.94 pp 714/1,962/89
16 s, no padding Lexical normalized 42.70% 33.22% +8.23 pp 702/1,963/86
16 s, Tafrigh noise Raw 57.08% 36.69% -0.22 pp 2,201/1,378/101
16 s, Tafrigh noise Leaderboard repo-exact 36.81% 25.31% -0.67 pp 876/1,387/110
16 s, Tafrigh noise Leaderboard intended 34.43% 24.75% -0.57 pp 726/1,381/110
16 s, Tafrigh noise Lexical normalized 33.93% 24.57% -0.54 pp 697/1,382/107
16 s, true silence Raw 55.61% 36.12% -1.69 pp 2,164/1,323/98
16 s, true silence Leaderboard repo-exact 36.76% 24.63% -0.71 pp 929/1,333/108
16 s, true silence Leaderboard intended 34.38% 24.10% -0.62 pp 779/1,326/109
16 s, true silence Lexical normalized 33.88% 23.93% -0.59 pp 750/1,327/106
Probe results by dataset and domain
Probe lexical WER by dataset
Dataset Configuration WER CER Δ Paired 95% CI Exact disagreements
Casablanca 15 s, Tafrigh noise (default) 50.09% 30.67% Baseline Not estimated 0 / 0.00%
Casablanca 15 s, true silence 48.59% 28.10% -1.50 pp [-3.67, +0.47] pp 60 / 60.00%
Casablanca 16 s, no padding 49.56% 28.38% -0.53 pp [-2.74, +1.25] pp 72 / 72.00%
Casablanca 16 s, Tafrigh noise 49.65% 30.02% -0.44 pp [-1.48, +0.50] pp 32 / 32.00%
Casablanca 16 s, true silence 48.42% 27.77% -1.67 pp [-3.41, -0.10] pp 57 / 57.00%
Common Voice 18 Arabic 15 s, Tafrigh noise (default) 12.30% 6.09% Baseline Not estimated 0 / 0.00%
Common Voice 18 Arabic 15 s, true silence 9.57% 3.02% -2.73 pp [-5.43, -0.20] pp 24 / 24.00%
Common Voice 18 Arabic 16 s, no padding 19.53% 12.37% +7.23 pp [+1.82, +13.00] pp 55 / 55.00%
Common Voice 18 Arabic 16 s, Tafrigh noise 11.52% 5.21% -0.78 pp [-2.52, +0.89] pp 16 / 16.00%
Common Voice 18 Arabic 16 s, true silence 9.57% 3.26% -2.73 pp [-5.15, -0.55] pp 21 / 21.00%
FLEURS ar-EG 15 s, Tafrigh noise (default) 26.44% 20.94% Baseline Not estimated 0 / 0.00%
FLEURS ar-EG 15 s, true silence 24.50% 18.57% -1.94 pp [-3.04, -0.96] pp 39 / 39.00%
FLEURS ar-EG 16 s, no padding 28.32% 22.51% +1.88 pp [+0.39, +3.40] pp 80 / 80.00%
FLEURS ar-EG 16 s, Tafrigh noise 26.44% 20.93% +0.00 pp [-0.68, +0.71] pp 26 / 26.00%
FLEURS ar-EG 16 s, true silence 24.30% 18.44% -2.14 pp [-3.46, -1.09] pp 40 / 40.00%
Quran ayah CA proxy 15 s, Tafrigh noise (default) 39.80% 36.13% Baseline Not estimated 0 / 0.00%
Quran ayah CA proxy 15 s, true silence 45.26% 41.16% +5.47 pp [+0.21, +11.29] pp 56 / 56.00%
Quran ayah CA proxy 16 s, no padding 73.32% 71.75% +33.53 pp [+7.90, +70.43] pp 79 / 79.00%
Quran ayah CA proxy 16 s, Tafrigh noise 37.97% 35.17% -1.82 pp [-5.47, -0.21] pp 37 / 37.00%
Quran ayah CA proxy 16 s, true silence 44.46% 40.80% +4.66 pp [-0.91, +12.54] pp 60 / 60.00%
SADA22 15 s, Tafrigh noise (default) 35.89% 22.07% Baseline Not estimated 0 / 0.00%
SADA22 15 s, true silence 33.70% 18.95% -2.19 pp [-3.80, -0.71] pp 62 / 62.00%
SADA22 16 s, no padding 36.03% 22.07% +0.14 pp [-1.88, +2.27] pp 68 / 68.00%
SADA22 16 s, Tafrigh noise 35.82% 22.28% -0.07 pp [-1.14, +1.04] pp 40 / 40.00%
SADA22 16 s, true silence 34.04% 19.90% -1.85 pp [-3.70, -0.26] pp 54 / 54.00%
Probe lexical WER by domain
Domain Configuration WER CER Δ Paired 95% CI Exact disagreements
Classical Arabic 15 s, Tafrigh noise (default) 39.80% 36.13% Baseline Not estimated 0 / 0.00%
Classical Arabic 15 s, true silence 45.26% 41.16% +5.47 pp [+0.21, +11.29] pp 56 / 56.00%
Classical Arabic 16 s, no padding 73.32% 71.75% +33.53 pp [+7.90, +70.43] pp 79 / 79.00%
Classical Arabic 16 s, Tafrigh noise 37.97% 35.17% -1.82 pp [-5.47, -0.21] pp 37 / 37.00%
Classical Arabic 16 s, true silence 44.46% 40.80% +4.66 pp [-0.91, +12.54] pp 60 / 60.00%
Dialect 15 s, Tafrigh noise (default) 42.10% 25.76% Baseline Not estimated 0 / 0.00%
Dialect 15 s, true silence 40.22% 22.88% -1.89 pp [-3.16, -0.70] pp 122 / 61.00%
Dialect 16 s, no padding 41.95% 24.78% -0.15 pp [-1.57, +1.28] pp 140 / 70.00%
Dialect 16 s, Tafrigh noise 41.87% 25.60% -0.23 pp [-0.97, +0.54] pp 72 / 36.00%
Dialect 16 s, true silence 40.33% 23.28% -1.77 pp [-3.08, -0.64] pp 111 / 55.50%
MSA 15 s, Tafrigh noise (default) 23.52% 18.20% Baseline Not estimated 0 / 0.00%
MSA 15 s, true silence 21.41% 15.70% -2.10 pp [-3.10, -1.13] pp 63 / 31.50%
MSA 16 s, no padding 26.51% 20.63% +2.99 pp [+1.38, +4.73] pp 135 / 67.50%
MSA 16 s, Tafrigh noise 23.35% 18.03% -0.16 pp [-0.79, +0.45] pp 42 / 21.00%
MSA 16 s, true silence 21.25% 15.64% -2.26 pp [-3.33, -1.22] pp 61 / 30.50%
Probe results for every dialect label
Probe lexical WER by dialect label
Dialect Configuration WER CER Δ Paired 95% CI Exact disagreements
Algeria 15 s, Tafrigh noise (default) 66.07% 33.57% Baseline Not estimated 0 / 0.00%
Algeria 15 s, true silence 59.82% 29.37% -6.25 pp [-12.87, +0.00] pp 7 / 50.00%
Algeria 16 s, no padding 63.39% 28.67% -2.68 pp [-10.59, +3.18] pp 10 / 71.43%
Algeria 16 s, Tafrigh noise 64.29% 31.47% -1.79 pp [-6.67, +2.22] pp 4 / 28.57%
Algeria 16 s, true silence 60.71% 29.90% -5.36 pp [-12.82, +1.96] pp 9 / 64.29%
Classical Arabic 15 s, Tafrigh noise (default) 39.80% 36.13% Baseline Not estimated 0 / 0.00%
Classical Arabic 15 s, true silence 45.26% 41.16% +5.47 pp [+0.21, +11.29] pp 56 / 56.00%
Classical Arabic 16 s, no padding 73.32% 71.75% +33.53 pp [+7.90, +70.43] pp 79 / 79.00%
Classical Arabic 16 s, Tafrigh noise 37.97% 35.17% -1.82 pp [-5.47, -0.21] pp 37 / 37.00%
Classical Arabic 16 s, true silence 44.46% 40.80% +4.66 pp [-0.91, +12.54] pp 60 / 60.00%
Egypt 15 s, Tafrigh noise (default) 30.58% 12.55% Baseline Not estimated 0 / 0.00%
Egypt 15 s, true silence 32.52% 14.20% +1.94 pp [-0.96, +6.17] pp 9 / 60.00%
Egypt 16 s, no padding 33.50% 13.17% +2.91 pp [-2.08, +7.78] pp 12 / 80.00%
Egypt 16 s, Tafrigh noise 30.10% 12.04% -0.49 pp [-1.59, +0.00] pp 4 / 26.67%
Egypt 16 s, true silence 31.55% 11.93% +0.97 pp [-1.40, +3.29] pp 8 / 53.33%
Egyptian 15 s, Tafrigh noise (default) 25.00% 5.56% Baseline Not estimated 0 / 0.00%
Egyptian 15 s, true silence 25.00% 8.33% +0.00 pp [+0.00, +0.00] pp 1 / 50.00%
Egyptian 16 s, no padding 25.00% 9.26% +0.00 pp [-7.69, +14.29] pp 2 / 100.00%
Egyptian 16 s, Tafrigh noise 25.00% 8.33% +0.00 pp [+0.00, +0.00] pp 1 / 50.00%
Egyptian 16 s, true silence 25.00% 8.33% +0.00 pp [+0.00, +0.00] pp 1 / 50.00%
Hijazi 15 s, Tafrigh noise (default) 29.85% 22.60% Baseline Not estimated 0 / 0.00%
Hijazi 15 s, true silence 21.64% 13.00% -8.21 pp [-15.18, -3.23] pp 10 / 58.82%
Hijazi 16 s, no padding 29.85% 21.83% +0.00 pp [-4.35, +4.40] pp 11 / 64.71%
Hijazi 16 s, Tafrigh noise 29.10% 20.74% -0.75 pp [-4.40, +2.91] pp 6 / 35.29%
Hijazi 16 s, true silence 23.88% 13.62% -5.97 pp [-13.27, -0.73] pp 8 / 47.06%
Jordan 15 s, Tafrigh noise (default) 29.27% 9.47% Baseline Not estimated 0 / 0.00%
Jordan 15 s, true silence 30.49% 8.16% +1.22 pp [-8.64, +14.93] pp 6 / 75.00%
Jordan 16 s, no padding 28.05% 8.42% -1.22 pp [-6.82, +4.82] pp 4 / 50.00%
Jordan 16 s, Tafrigh noise 30.49% 10.53% +1.22 pp [+0.00, +3.57] pp 4 / 50.00%
Jordan 16 s, true silence 28.05% 7.37% -1.22 pp [-9.38, +9.59] pp 6 / 75.00%
Khaliji 15 s, Tafrigh noise (default) 52.17% 37.43% Baseline Not estimated 0 / 0.00%
Khaliji 15 s, true silence 46.09% 30.20% -6.09 pp [-15.46, +0.49] pp 11 / 57.89%
Khaliji 16 s, no padding 51.30% 38.88% -0.87 pp [-7.50, +4.17] pp 10 / 52.63%
Khaliji 16 s, Tafrigh noise 54.78% 41.59% +2.61 pp [-1.77, +8.18] pp 5 / 26.32%
Khaliji 16 s, true silence 49.57% 33.82% -2.61 pp [-8.97, +3.38] pp 8 / 42.11%
MSA 15 s, Tafrigh noise (default) 23.52% 18.20% Baseline Not estimated 0 / 0.00%
MSA 15 s, true silence 21.41% 15.70% -2.10 pp [-3.05, -1.21] pp 63 / 31.50%
MSA 16 s, no padding 26.51% 20.63% +2.99 pp [+1.34, +4.61] pp 135 / 67.50%
MSA 16 s, Tafrigh noise 23.35% 18.03% -0.16 pp [-0.83, +0.48] pp 42 / 21.00%
MSA 16 s, true silence 21.25% 15.64% -2.26 pp [-3.37, -1.29] pp 61 / 30.50%
Mauritania 15 s, Tafrigh noise (default) 89.95% 83.41% Baseline Not estimated 0 / 0.00%
Mauritania 15 s, true silence 92.06% 85.70% +2.12 pp [+0.00, +5.14] pp 5 / 38.46%
Mauritania 16 s, no padding 87.83% 79.98% -2.12 pp [-8.29, +1.82] pp 7 / 53.85%
Mauritania 16 s, Tafrigh noise 88.89% 81.58% -1.06 pp [-3.68, +1.23] pp 3 / 23.08%
Mauritania 16 s, true silence 88.89% 83.41% -1.06 pp [-4.80, +1.58] pp 4 / 30.77%
More than 1 speaker اكثر من متحدث 15 s, Tafrigh noise (default) 35.12% 21.35% Baseline Not estimated 0 / 0.00%
More than 1 speaker اكثر من متحدث 15 s, true silence 34.33% 19.88% -0.78 pp [-3.05, +1.18] pp 16 / 76.19%
More than 1 speaker اكثر من متحدث 16 s, no padding 34.60% 20.44% -0.52 pp [-3.44, +2.62] pp 19 / 90.48%
More than 1 speaker اكثر من متحدث 16 s, Tafrigh noise 34.99% 21.86% -0.13 pp [-1.80, +1.66] pp 17 / 80.95%
More than 1 speaker اكثر من متحدث 16 s, true silence 34.07% 21.14% -1.04 pp [-3.64, +1.14] pp 15 / 71.43%
Morocco 15 s, Tafrigh noise (default) 47.49% 23.43% Baseline Not estimated 0 / 0.00%
Morocco 15 s, true silence 39.66% 14.08% -7.82 pp [-16.32, -0.56] pp 12 / 70.59%
Morocco 16 s, no padding 45.25% 16.45% -2.23 pp [-10.06, +3.43] pp 13 / 76.47%
Morocco 16 s, Tafrigh noise 48.60% 23.91% +1.12 pp [-1.90, +3.89] pp 6 / 35.29%
Morocco 16 s, true silence 45.25% 19.64% -2.23 pp [-7.65, +1.96] pp 9 / 52.94%
Najdi 15 s, Tafrigh noise (default) 26.57% 12.48% Baseline Not estimated 0 / 0.00%
Najdi 15 s, true silence 25.37% 11.02% -1.19 pp [-3.50, +0.73] pp 17 / 77.27%
Najdi 16 s, no padding 26.87% 12.61% +0.30 pp [-2.49, +3.17] pp 17 / 77.27%
Najdi 16 s, Tafrigh noise 26.57% 12.24% +0.00 pp [-1.11, +0.87] pp 8 / 36.36%
Najdi 16 s, true silence 26.27% 11.02% -0.30 pp [-2.50, +1.56] pp 14 / 63.64%
Notapplicable 15 s, Tafrigh noise (default) 70.00% 57.45% Baseline Not estimated 0 / 0.00%
Notapplicable 15 s, true silence 60.00% 34.04% -10.00 pp [-20.00, +0.00] pp 1 / 33.33%
Notapplicable 16 s, no padding 70.00% 57.45% +0.00 pp [+0.00, +0.00] pp 0 / 0.00%
Notapplicable 16 s, Tafrigh noise 50.00% 46.81% -20.00 pp [-100.00, +0.00] pp 1 / 33.33%
Notapplicable 16 s, true silence 60.00% 34.04% -10.00 pp [-20.00, +0.00] pp 1 / 33.33%
Palestine 15 s, Tafrigh noise (default) 28.36% 10.03% Baseline Not estimated 0 / 0.00%
Palestine 15 s, true silence 34.33% 10.64% +5.97 pp [+1.20, +13.64] pp 5 / 50.00%
Palestine 16 s, no padding 34.33% 16.72% +5.97 pp [+0.00, +12.68] pp 6 / 60.00%
Palestine 16 s, Tafrigh noise 29.85% 9.42% +1.49 pp [+0.00, +6.52] pp 4 / 40.00%
Palestine 16 s, true silence 29.85% 9.42% +1.49 pp [+0.00, +6.52] pp 4 / 40.00%
UAE 15 s, Tafrigh noise (default) 50.00% 29.55% Baseline Not estimated 0 / 0.00%
UAE 15 s, true silence 44.34% 22.11% -5.66 pp [-11.24, +0.00] pp 10 / 83.33%
UAE 16 s, no padding 46.23% 23.97% -3.77 pp [-7.29, +0.00] pp 11 / 91.67%
UAE 16 s, Tafrigh noise 50.00% 30.37% +0.00 pp [+0.00, +0.00] pp 2 / 16.67%
UAE 16 s, true silence 45.28% 22.31% -4.72 pp [-10.23, +0.00] pp 9 / 75.00%
Unknown 15 s, Tafrigh noise (default) 67.50% 46.82% Baseline Not estimated 0 / 0.00%
Unknown 15 s, true silence 63.75% 38.17% -3.75 pp [-8.97, +1.64] pp 6 / 37.50%
Unknown 16 s, no padding 75.00% 53.18% +7.50 pp [-4.17, +20.51] pp 9 / 56.25%
Unknown 16 s, Tafrigh noise 67.50% 44.53% +0.00 pp [+0.00, +0.00] pp 2 / 12.50%
Unknown 16 s, true silence 60.00% 37.40% -7.50 pp [-15.19, +0.00] pp 7 / 43.75%
Yemen 15 s, Tafrigh noise (default) 41.54% 21.26% Baseline Not estimated 0 / 0.00%
Yemen 15 s, true silence 40.00% 17.78% -1.54 pp [-6.71, +2.72] pp 6 / 54.55%
Yemen 16 s, no padding 41.54% 20.94% +0.00 pp [-3.79, +2.30] pp 9 / 81.82%
Yemen 16 s, Tafrigh noise 39.49% 19.96% -2.05 pp [-5.62, +0.00] pp 5 / 45.45%
Yemen 16 s, true silence 39.49% 15.59% -2.05 pp [-7.48, +1.46] pp 8 / 72.73%
Probe decision: no-padding was decisively worse overall (+8.23 pp, 95% CI [+1.13, +18.40]) and catastrophic for the Classical Arabic sample. The padded alternatives disagreed frequently and no padded candidate was a statistically decisive overall winner. The full run therefore retained exact Tafrigh default behavior: 15-second cuts with one second of generated audio on each side.

Tafrigh calls WhiteNoise(..., volume=0). In Pydub, 0 means 0 dB gain rather than zero amplitude; measured output was roughly −4.8 dBFS. “True silence” used −120 dB. Wit outputs were not deterministic across calls/apps, including one minor Quran variant on the same encoded payload.

Wit/Tafrigh pipeline and telemetry

Full Wit run telemetry
Measure Value Meaning
Successful speech-segment requests 32,160 HTTP 200 segment recognitions
Recorded HTTP attempts 32,192 Includes retries
Recovered retries 32 0.099% of attempts
Unresolved failures 0 No transport/API failure scored as empty text
No-speech clips 82 Auditok detected no region
Successful empty segments 4,980 HTTP 200 with empty text
Clips with one or more empty segments 3,954 Still checkpointed as successful
Detected speech 31.731 h 87.19% of decoded input
Whole-source WAV→MP3 conversions 13,943 Tafrigh compatibility preprocessing
MP3 request payload 1.793 GB Aggregate uploaded bytes
Forced MP3 demux fallbacks 1 One valid file was mis-probed by Auditok/FFmpeg
Preprocessing retries 0 None required

Compatibility boundary

  1. Non-MP3 clips are converted in full to MP3 with Pydub/FFmpeg.
  2. Tafrigh's Auditok splitter uses min_dur=0.5, max_silence=0.5, energy_threshold=50, and a 15-second maximum region.
  3. Each region is re-encoded after one second of Tafrigh-generated noise is added at both sides.
  4. Requests use POST /speech, audio/mpeg3, Wit media version 20200513, Arabic question-mark substitution, and chronological segment concatenation.

Scheduler

Stock Tafrigh loops over files sequentially and creates a manager/process pool for every file. On 24,414 short clips, even a one-request-per-file admission lower bound is 6h 47m. The benchmark preserves Tafrigh's recognition inputs but shares a bounded scheduler across files, uses 8 distinct Arabic apps in 2 inferred user groups, limits each app to one start every 1.02s, caps each user group at 4 starts/s, and checkpoints a clip only after every segment succeeds. Every retry reacquires both quota limits.

The operator's production topology is eight Tafrigh processes with eight keys each, and the 64 keys have independent app/user quotas. That can materially outperform a single eight-app run on a folder when enough independent files exist: the processes can upload and recognize separate regions concurrently. It does not make one unsharded file automatically eight times faster, and actual aggregate scaling still depends on Auditok/encoding CPU, upload bandwidth, network latency, API service time, and per-process scheduling. Cohere has the opposite shape: one GPU model is best kept resident and globally batched across files rather than duplicated eight times on a 12 GB card.

Methodology and reproducibility

Evaluation method
Item Specification
Primary engines Cohere Arabic Transcribe model; Wit.ai /speech through Tafrigh 1.7.8-compatible preprocessing
Cohere revision 0a8193caa4f3f92131471ab08824e488141cb392
July 12 production script d1dda35f9d393ee11001d4fe71c7c6db255f76c4320f90baf354429d7423b510
Production runtime bound Transformers >=5.13,<5.14; portable dependencies in final/requirements.txt
Production GPU NVIDIA GeForce RTX 3060, 11.754 GiB, compute capability 8.6
Tafrigh revision 2ccba42db8c34c04924d1befc35cb3d5eec80d93
Native runtime CrispASR/GGML CUDA research branch; GGML 0714117daca2471b00e09554c7eaa74a06b0b2c5
Native GGUF gate 500 clips / 5,032.699 s; E4/D8; greedy device argmax; F16, Q8_0, Q4_K imatrix
Sample pairing 24,414 matching IDs, references, metadata records, and audio paths
Timing denominator 131,015.81575 decoded seconds at 16 kHz; 36.393282 hours
Primary statistic Corpus WER: total S+D+I divided by total reference words
Primary profile Lexical normalized
Uncertainty 5,000-replicate paired cluster percentile bootstrap; 95% interval; seed 0
Cluster rule Speaker within dataset where present; otherwise utterance
Overall clusters 14,319 (976 speaker clusters + 13,343 utterance clusters)
Probe uncertainty 2,000 paired cluster-bootstrap replicates; same confidence and seed
Native quantization uncertainty 20,000 paired utterance-bootstrap replicates; 95% interval; seed 20260711
WER tie-breaking NeMo/Kaldi-compatible insertion, then deletion, then substitution on equal-cost paths

Validation gates

Environment

Wit used Python 3.12.12, tafrigh[wit]==1.7.8, and Requests 2.34.2 in an isolated environment without PyTorch, Transformers, Whisper, or Cohere dependencies. The final production profile used Python 3.12.12, PyTorch/TorchAudio 2.11.0+cu128, Transformers 5.13.0, ONNX Runtime 1.27.0, and CUDA BF16 on an NVIDIA GeForce RTX 3060. Exact tested package versions remain recorded in the production profile and release summary; the portable bundle intentionally constrains only compatibility-sensitive packages and leaves device-specific Torch wheels to the installer.

Limitations and correct interpretation

Sources, versions, and artifact integrity

Public references

Machine-readable source artifacts

Source artifacts and SHA-256 checksums
Artifact SHA-256 Size/scope
benchmark/reports/final_cohere_vs_wit_20260711.json c8306bcd7f70f006e271950a0cf4a052741cfed5e316ffb0a90ebc401f91a12f 187,895 bytes
benchmark/reports/wit_probe500.json bd73f5668a3e7874038f00dd83949e9e3d3cd6a0190db92cba37425c767c67ab 439,824 bytes
benchmark/reports/full_exact_20260710.json ea39c3ea18b8c1527c37d03fcebbb71b5aa451fd79e633ae8a3d422b943253d8 385,534 bytes
benchmark/reports/final_full_20260711.json bd49322844748b87d34679ac716eaefa4c0694db637b631c0d40253657ad66be 186,000 bytes
benchmark/reports/pipeline_probe500_20260711.json d77464eaebc2d8b22a6471e77b081d333a53756cf5688664fd67a0c622f8fdda 168,734 bytes
benchmark/results/wit_full_default_20260710/summary.json 2726e15b57156298849b5315d9fff6bbca471a1983924afda653e1209d3f3c89 4,380 bytes
benchmark/results/longform_precision_probe.json 0b7617a54e769009b7ea7b92242640bf8c6e01335dcbd66c98b35298b31216e7 4,025 bytes
benchmark/reports/timestamp_modes_500_20260711.json 8947fb3ae5d8023d73114a66e88586143262e58eed773a6e00ce75394db094d1 444,127 bytes
benchmark/results/timestamp_modes_20260711/performance.json 590b6cc16c9e49f1b13da1dd8089b9fd065d46fd56fb7e1a96a0b205963d0bd4 3,375 bytes
benchmark/results/alignment_modes_frozen_20260711/summary.json ddab3281d0e0d552daef924a2970fa535ee54b30126a9bd8baff97a2f7679947 14,855 bytes
benchmark/results/longform_timestamp_frozen_20260711/drift.json 479dcbdab2d5ed3e3ef769f2c2a71752b2aee2fe73cf6a1d4e58e7903f30aede 149,929 bytes
benchmark/results/longform_text_only_20260711/summary.json c5b883312e65be4f1448a06fcd4fb555d1bf6b1f77f3ea0041ae0537f54eb832 692 bytes
benchmark/reports/vad_modes_500_20260711.json 1c1e825dd1ca688a4badbc6d7efb5bef10f8e3e10de128ebe28427bb31220237 30,682 bytes
benchmark/reports/vad_modes_500_20260711.md c1ecf1f7938ba42778f362811adc744aa25d61d3a7afcc1b8611b3ce60c50ed4 2,658 bytes
benchmark/results/vad_modes_20260711/segmentation_audit.json b587656a53a811218e69fff6364486069f4a9d6f8a27af1fad6d99cb8b32fbeb 4,241 bytes
benchmark/results/vad_modes_20260711/segmentation_audit.time.txt 50b9c9caf53be327ecbb80133ccb6f9418a4e97dcf0c0f9140c95a049fddde8e 60 bytes
benchmark/reports/production_release_20260712.json 74f4caf4f49691357ab160711cf07b811d959b7443c2aa470e50d29768e87e2a 34,982 bytes
benchmark/PRODUCTION_REVIEW_20260712.md 00140f207a86b06c01b2630b33896ed9e0915a9f99df34512f36ed4311f43976 9,584 bytes
benchmark/reports/cohere_wit_playback_comparison_20260712.html c0b5250163dae64345afb21ca725b893f8629b17e892d9b5b615edee833f62e3 233,834 bytes
benchmark/reports/cohere_wit_playback_comparison_20260712.metadata.json fcd664aaaf0269fbfcb5d677b085aca3682e66b2dd8ba171d97c7ece9f332f16 5,172 bytes
benchmark/reports/cohere_wit_playback_comparison_20260712.pastehtml.json ed841063514d088824202352dfe9833c02a7961169a7b2e142f38857e5b147e7 654 bytes
benchmark/reports/build_cohere_wit_playback_demo.py edc3b73a1cbd32ee7b11bfef15bac5e3b2ae08b8bac4de52755f96cac089b28b 34,993 bytes
benchmark/reports/cohere_wit_comparison_assets/1-original.json 99fa81eed7469a06c03c1aa4bb5d6b72c1f80228c1426c5bf818d69dd6816055 117,998 bytes
benchmark/reports/cohere_wit_comparison_assets/1.json 28dac6cb48cf4c3607e0be388bebee2fbd347c49880a4616253514c8785d0d49 111,880 bytes
native_benchmark/results/quantized_full500_20260711/comparison.json 2cba3089d76024fd469203f2765a7ad431ac63faf1ec55fc41201727f8b7563f 417,774 bytes
native_benchmark/results/quantized_full500_20260711/REPORT.md 2e52d1c589c9de2268ef5558fbb448b6c43be284bbd96e1d50baf4488b7eb2a4 5,786 bytes
native_benchmark/results/frontend_parallel/benchmark.json 17a5a5b4a5b9999ec5022e2dc582be2daeeaa81c30066ef8ff92d6818f8bbc29 2,124 bytes
native_benchmark/DECODER_BATCH_PARITY.md 0980c98a065a0382c16a7c1073c02b3d65c03f569ec17381202db00e2a9331e8 10,417 bytes
native_benchmark/results/multistream_reprobe_20260711/summary.json 3e6f632c0b4aae989611e32e9852e7f1c9ed7672f0e45991fda00a1770b55ffe 1,609 bytes
native_benchmark/results/multistream_reprobe_20260711/REPORT.md 86a0136352cebb8c6d7b16cbc9e0ea4bc92a2fdcb1cb04c2cf12c3121bbd2a66 4,706 bytes
native_benchmark/results/native_onewav_release_20260711/summary.json 1c28f4e6f7aacba7075e419b088692e3a28cb99717b6b6fa7ed112ccb6e5a27f 23,362 bytes
native_benchmark/results/native_onewav_release_20260711/REPORT.md b0b75971004f549ca4a1531db31a4d3b5bfdcfaedc0d0fc200a70a7cd3b5b7c3 1,911 bytes
benchmark/UPSTREAM_CANDIDATES_20260712.md d84275834d55bfa4a0a6826b92f8c6d7e13f5af356b8e01059da6bcd34c975f3 27,209 bytes
final/transcribe.py e27add42e280781e0a98efa1e032a2675fcd9feb459d5e8a2e736e01e1d34e78 164,834 bytes
final/README.md 82983546fa37028bb5a95a398d27d34c8809dde509a61f33451658c98ac649f9 1,864 bytes
final/VERSION.json 8f4a6215c32133deab8dc5cb015c99e46fd459c38029405cef956191dbff4d96 921 bytes
final/requirements.txt 19ecd88291e7bf3c70f43b3761a760470c0aedf2b0a7e9b61cc97d617349ebfb 982 bytes
final/requirements-optional.txt 33680ee16a154dd87de2c52f06082e7aa7c9472f8d8aa67e9f200063eb7ba292 345 bytes
final/SETUP.md 90b78c77fef726aa08132738874889acfd19bb8a617062c7af6f7d43a75fa2a9 10,089 bytes
final/validate_install.py a521b02be15354b3f479b0f4889ed85c76994db6edf9e0d6946de1fe2ed71699 10,719 bytes
final/transcribe_assets/README.md 3003b2763c12abd0f1d497cded37bf91c6d9094bb6961ff25b15662e17bf2509 1,216 bytes
final/transcribe_assets/LICENSE.silero-vad 51c19c8be941a3fb00ccf58f0bf9053de9f7237a0b37327896eabad32dffe873 1,076 bytes
final/transcribe_assets/LICENSE.faster-whisper af6798135e729f8aa6c853936d037dfdea449734d26b8ea6a89805fca758c0d5 1,064 bytes
benchmark/manifests/wer_eval.jsonl 332aa2a063cf285a584bc1630f9164c15d18ea904b79039c39585cf9919a277d 24,414 JSONL rows

This HTML is a derived, static rendering. Percentages are rounded for display; conclusions and intervals come from the unrounded JSON values. Local absolute paths, credentials, transcript text, and app identifiers are intentionally excluded.