PTG Fleet Model Leaderboards
Open-weight models benchmarked on Petronella Technology Group's private GPU fleet
(6x H200 NVL, RTX PRO 6000 Blackwell, GB10). Every published score is cross-judged: the generator
never grades itself, and the campaign judge is locked to GPT-4.1 (Session-13/15 lock-in; Haiku-4.5
or validated-equivalent DeepSeek-V4-Flash on legacy rows). Last updated: 2026-08-15 12:45 EDT.
Auto-generated from harness result files; the per-row Measured column is the data-collection date.
Current picks by hardware role (verdict 2026-07-05; re-verified 2026-07-06 on the v2 hard set)
4-way H200 NVL bridge (575 GB)
Ornith-1.0-397B-FP8 (MIT):
leads the v2 hard set outright (0.971, n=3, 20/20 every rep) and scored a perfect blog 1.000 on
all nine runs across three topics; research 0.91 at 120 t/s single-stream. Alternatives:
MiniMax-M3-MXFP8 (v2 0.931; 1M ctx, multimodal; restrictive community license) and GLM-5.2-NVFP4
(v2 0.942 but 150 s per case with high variance; still the research leader at 0.94).
2-way H200 NVL bridge (287 GB)
Qwen3-Coder-Next-FP8
(80B-A3B, Apache-2.0): best coder in the 2-way class on the v2 hard set (0.939, n=3, at 2.6 s per
case; the nearest class rival needs 20x the latency for less), 85.2% on the 88-case tool corpus,
178 t/s single-stream and 2,789 t/s aggregate at c=32. Long-form writer alternative:
MiniMax-M2.7-NVFP4 (blog 0.94 mean across three topics, v2 0.901).
Single 96 GB card
Qwen3.6-35B-A3B-FP8: best all-rounder
that fits one card (coding 0.975, wins research and cited-RAG); no displacer found (2026-05-31).
Long context (524K)
DeepSeek-V4-Flash: coding 0.964,
blog 0.944, SWE-bench Verified 67.3% measured locally (2026-06-28); also the fleet's validated
free grader.
Voice tool-caller
gpt-oss-20b: 125/125 correct declines,
0% tool hallucination on the 88-case corpus (2026-05-29).
Air-gap edge appliance
Ministral-3-8B (2026-05-27 re-test);
Granite-4.1-8B and Gemma-4-e4b for tool-driving. All Mistral models need vLLM
--tool-call-parser mistral.
Coding v1 (22 PTG ops/SEO/debug tasks; longitudinal reference - saturated at the top)
| Model | Score | Runs (range) | Pass | Latency | Temp | Measured |
| Claude Opus-4.8 (cloud reference) BEST | 0.995 | n=1 | 22/22 | 2.6s | default | 2026-07-06 |
| Claude Opus-4.7 | 0.991 | n=1 | 22/22 | 2.7s | default | 2026-05-25 |
| Laguna-S-2.1-NVFP4 | 0.991 | n=1 | 22/22 | 3.3s | default | 2026-07-27 |
| Qwen3.6-27B BF16 (dense) | 0.990 | n=1 | 20/22 | 66.4s | default | 2026-06-08 |
| gpt-oss:20b | 0.989 | n=1 | 22/22 | 2.5s | default | 2026-05-28 |
| DeepSeek-V4-Flash-DSpark (2x GB10) | 0.986 | n=1 | 22/22 | 11.2s | default | 2026-06-30 |
| Qwen3.6-27B (dense) | 0.984 | n=1 | 22/22 | 25.2s | default | 2026-06-30 |
| MiniMax-M3 (MXFP8) | 0.984 | n=3 (0.975-0.991) | 21/22 | 5.2s | default | 2026-07-05 |
| Laguna-M.1-NVFP4 | 0.984 | n=1 | 22/22 | 6.4s | default | 2026-07-26 |
| DeepSeek-V4-Pro (cloud reference) | 0.982 | n=1 | 22/22 | 4.1s | default | 2026-06-28 |
| GPT-5.2-Codex (cloud reference) | 0.982 | n=1 | 22/22 | 2.6s | default | 2026-07-06 |
| Ornith-1.0-397B (FP8) | 0.980 | n=3 (0.977-0.982) | 22/22 | 2.7s | default | 2026-07-05 |
| Qwen3.6-35B-A3B-NVFP4-Fast (unsloth) | 0.980 | n=5 (0.971-0.986) | 22/22 | 20.0s | default | 2026-07-12 |
| Gemma-4-31B | 0.980 | n=1 | 22/22 | 28.9s | default | 2026-05-27 |
| Nex-N2-Pro (NVFP4) | 0.980 | n=1 | 22/22 | 18.1s | default | 2026-06-10 |
| North-Mini-Code-1.0 (FP8, Cohere) | 0.980 | n=5 (0.976-0.986) | 22/22 | 3.4s | default | 2026-06-20 |
| Qwen3-Coder-Next-80B | 0.977 | n=3 (0.977-0.977) | 22/22 | 0.6s | default | 2026-07-05 |
| Gemma-4-12B | 0.977 | n=1 | 22/22 | 15.8s | default | 2026-06-03 |
| Qwen3.8-27B (BF16) | 0.977 | n=1 | 22/22 | 35.7s | default | 2026-08-14 |
| MiniMax-M2.7 (NVFP4) | 0.977 | n=3 (0.965-0.984) | 22/22 | 17.3s | default | 2026-07-05 |
| gemma-4-26b-a4b-nvfp4 | 0.977 | n=3 (0.970-0.980) | 22/22 | 1.5s | default | 2026-07-15 |
| Qwen3.6-35B-A3B | 0.975 | n=5 (0.961-0.984) | 22/22 | 0.6s | default | 2026-05-29 |
| gpt-oss-20b | 0.975 | n=1 | 22/22 | 11.8s | default | 2026-05-28 |
| ornith-35b | 0.975 | n=1 | 4/22 | 6.6s | default | 2026-06-29 |
| glm-5.2 | 0.975 | n=1 | 22/22 | 6.3s | default | 2026-07-13 |
| GLM-4.7-Flash | 0.973 | n=1 | 22/22 | 8.9s | default | 2026-05-28 |
| Qwen3.6-35B-A3B (BF16) | 0.973 | n=1 | 22/22 | 22.7s | default | 2026-06-08 |
| GLM-5.2 (NVFP4) | 0.972 | n=5 (0.959-0.986) | 22/22 | 46.0s | default | 2026-07-03 |
| Ornith-1.0-397B (W4A16 int4) | 0.970 | n=5 (0.959-0.982) | 21/22 | 3.6s | default | 2026-07-03 |
| Gemma-4-26B-A4B (BF16) | 0.970 | n=1 | 22/22 | 1.0s | default | 2026-06-09 |
| GPT-5.4 (cloud reference) | 0.970 | n=1 | 22/22 | 1.7s | default | 2026-07-06 |
| Laguna-XS-2.1-NVFP4 | 0.970 | n=1 | 21/22 | 0.7s | default | 2026-07-26 |
| Mistral-Medium-3.5-128B | 0.968 | n=1 | 22/22 | 4.6s | default | 2026-05-26 |
| Qwen3.6-27B-NVFP4 (unsloth, dense) | 0.967 | n=5 (0.966-0.970) | 21/22 | 14.4s | default | 2026-07-12 |
| qwen3-coder:480b | 0.966 | n=1 | 22/22 | 2.1s | default | 2026-06-28 |
| Qwen3.8-27B (FP8) | 0.966 | n=1 | 21/22 | 124.6s | default | 2026-08-14 |
| gpt-oss-120b | 0.964 | n=1 | 21/22 | 8.4s | default | 2026-05-28 |
| DeepSeek-V4-Flash | 0.964 | n=5 (0.961-0.970) | 21/22 | 0.9s | default | 2026-05-29 |
| GLM-5.2 (IQ2_M 2-bit) | 0.959 | n=1 | 21/22 | 26.9s | default | 2026-06-20 |
| nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.958 | n=1 | 21/22 | 18.8s | default | 2026-08-11 |
| gpt-4.1 | 0.957 | n=1 | 21/22 | 1.4s | default | 2026-05-25 |
| Mistral-Small-4-119B | 0.957 | n=1 | 21/22 | 1.4s | default | 2026-05-27 |
| gemma4-coder | 0.957 | n=1 | 22/22 | 3.1s | default | 2026-07-14 |
| GLM-4.5-Air | 0.956 | n=1 | 17/22 | 105.3s | default | 2026-05-28 |
| claude-haiku-4-5-20251001 | 0.956 | n=5 (0.955-0.959) | 21/22 | 1.2s | default | 2026-05-29 |
| Gemma-4-31B (BF16) | 0.955 | n=1 | 21/22 | 4.9s | default | 2026-06-09 |
| z-ai/glm-5.2 | 0.955 | n=1 | 21/22 | 17.3s | default | 2026-06-27 |
| Qwen3-Coder-30B | 0.952 | n=1 | 22/22 | 0.8s | default | 2026-06-09 |
| openai/gpt-oss-20b | 0.952 | n=1 | 21/22 | 1.8s | default | 2026-05-26 |
| qwen3-coder:30b | 0.952 | n=1 | 22/22 | 0.8s | default | 2026-05-27 |
| Gemma-4-12B (BF16) | 0.952 | n=1 | 21/22 | 2.5s | default | 2026-06-09 |
| Gemma-4-e4b (BF16) | 0.949 | n=1 | 22/22 | 1.8s | default | 2026-06-09 |
| Muse-Glimmer-30B (Meta) | 0.949 | n=1 | 21/22 | 14.6s | default | 2026-08-15 |
| Ornith-1.0-35B (FP8) | 0.948 | n=5 (0.913-0.970) | 21/22 | 3.7s | default | 2026-07-02 |
| Qwen3.8-27B (NInfer int4+MTP3) | 0.941 | n=1 | 21/22 | 18.8s | default | 2026-08-15 |
| command-a-plus | 0.939 | n=1 | 21/22 | 18.1s | default | 2026-05-26 |
| Gemma-4-e2b | 0.935 | n=1 | 21/22 | 12.0s | default | 2026-05-26 |
| Qwen3.5-Opus-distill (27B) | 0.934 | n=1 | 21/22 | 13.1s | default | 2026-05-26 |
| Qwen3.8-27B (Q4_K_XL) | 0.934 | n=1 | 20/22 | 171.1s | default | 2026-08-15 |
| Devstral-Small-2-24B | 0.932 | n=1 | 20/22 | 6.7s | default | 2026-05-28 |
| gemma-4-12b-nvfp4 | 0.931 | n=3 (0.930-0.932) | 21/22 | 2.1s | default | 2026-07-15 |
| Gemma-4-e4b | 0.931 | n=1 | 20/22 | 13.6s | default | 2026-05-26 |
| Falcon-H1R-7B | 0.926 | n=1 | 19/22 | 84.2s | default | 2026-05-27 |
| devstral-small-2-24b | 0.925 | n=5 (0.900-0.936) | 20/22 | 0.7s | default | 2026-07-03 |
| Mixtral-8x22B | 0.909 | n=1 | 21/22 | 1.9s | default | 2026-05-26 |
| Mistral-Small-24B | 0.909 | n=1 | 20/22 | 1.4s | default | 2026-06-09 |
| Granite-4.1-8B | 0.903 | n=1 | 19/22 | 0.8s | default | 2026-06-09 |
| hf.co/tiiuae/Falcon-H1R-7B-GGUF:Q4_K_M | 0.885 | n=1 | 19/22 | 49.9s | default | 2026-05-27 |
| q25-gptq | 0.885 | n=1 | 19/22 | 2.4s | default | 2026-05-29 |
| Apriel-1.6-15B | 0.880 | n=1 | 19/22 | 110.6s | default | 2026-05-27 |
| Gemma-4-26B-A4B | 0.877 | n=1 | 18/22 | 18.4s | default | 2026-05-26 |
| ornith-9b | 0.875 | n=1 | 19/22 | 54.0s | default | 2026-06-29 |
| Ministral-3-8B | 0.872 | n=1 | 18/22 | 2.9s | default | 2026-05-27 |
| hf.co/unsloth/Ministral-3-8B-Instruct-2512-GGUF:Q4_K_M | 0.858 | n=1 | 18/22 | 0.8s | default | 2026-05-27 |
| qwen2.5:7b-instruct | 0.847 | n=1 | 18/22 | 2.1s | default | 2026-05-29 |
| l31-awq | 0.847 | n=1 | 19/22 | 3.4s | default | 2026-05-29 |
| LFM2.5-8B-A1B (Liquid) | 0.846 | n=1 | 17/22 | 2.0s | default | 2026-06-03 |
| q25-awq | 0.820 | n=1 | 17/22 | 2.5s | default | 2026-05-29 |
| llama3.1:8b | 0.820 | n=1 | 17/22 | 2.7s | default | 2026-05-29 |
| l31-gptq | 0.794 | n=1 | 15/22 | 2.9s | default | 2026-05-29 |
| Mixtral-8x7B | 0.773 | n=1 | 17/22 | 0.7s | default | 2026-05-26 |
| MiniCPM5-1B | 0.644 | n=1 | 10/22 | 6.2s | default | 2026-05-27 |
Incumbent fleet coder
Qwen3.6-35B-A3B now
co-leads.
North-Mini-Code-1.0
(Cohere, Apache-2.0, 30B-total/3B-active MoE) matches the very top on PTG coding -
0.980 ± 0.004 (N=5),
statistically tied with Nex-N2-Pro and edging Qwen3.6-35B-FP8 (0.975) -
and loop-amplifies on agentic
tasks (cap=1 0.833 → cap=5 0.917, fabrication 2→1), unlike the non-amplifiers below. Apache-2.0, FP8 fits one card,
fast, clean tools; a genuine alternative fleet coder (Session-22). It is not a research/RAG model (0.821) and blogs short of 3,000 words.
Gemma-4-12B (0.977) ties the 31B/Qwen3.6 on single-shot coding but is single-shot only -
it does NOT loop-amplify (cap=1=cap=5=0.917 on agentic tasks) and fabricates "done" under loop pressure;
use it as a fast coder, not an agent.
LFM2.5-8B-A1B is an on-device model (edge-tier 0.846, 3.41%
tool-call hallucination) - not a fleet upgrade.
Score = mean of each model's largest rep family (newest on ties); scores within 0.03 of each other are statistical ties. Claude Opus-4.7 reference = 1.00. Scores cross-judged by GPT-4.1, Haiku-4.5, or validated-equivalent
DeepSeek-V4-Flash. Temperature column: stamped value where present;
"default" = pre-Session-13 runs (eval_coding.py default 0.0).
Per-case breakdown: coding_tasks.jsonl (22 cases; full prompt texts withheld to keep the benchmark uncontaminated)
| Case id | Category | Automatic checks | Judge scale |
|---|
| 01_bash_backup_recent_html | bash_ops | syntax:bash, must_include(3), must_include_any(3) | 1-5 |
| 02_fish_fleet_uptime | bash_ops | syntax:fish, must_include(3), must_include_any(2) | 1-5 |
| 03_bash_safe_rsync_deploy | bash_ops | syntax:bash, must_include(3), must_include_any(2) | 1-5 |
| 04_py_openai_compatible_call | python_automation | syntax:python, must_include(3), must_not_include(2) | 1-5 |
| 05_py_factorial_constraints | instruction_following | syntax:python, must_include(1), must_not_include(2) | 1-5 |
| 06_py_retry_backoff_decorator | python_automation | syntax:python, must_include(2), must_include_any(3) | 1-5 |
| 07_py_concurrent_endpoint_ping | python_automation | syntax:python, must_include(2), must_include_any(3) | 1-5 |
| 08_py_parse_nvidia_smi_power | python_automation | syntax:python, must_include(3) | 1-5 |
| 09_py_phone_regex | python_automation | syntax:python, must_include(2), must_include_any(2) | 1-5 |
| 10_systemd_timer_oncalendar | config_edit | must_include(3), must_not_include(1) | 1-5 |
| 11_yaml_add_fleet_tier | config_edit | must_include(5) | 1-5 |
| 12_htaccess_301_https | config_edit | must_include(4), must_include_any(3) | 1-5 |
| 13_php_canonical_tag | python_automation | must_include(2), must_include_any(2) | 1-5 |
| 14_sql_top_referring_domains | sql | must_include(5), must_include_any(4) | 1-5 |
| 15_debug_keyerror | debug_fix | syntax:python, must_include(1), must_include_any(4) | 1-5 |
| 16_debug_ollama_cpu_only | debug_fix | must_include_any(6) | 1-5 |
| 17_refactor_bash_loop | refactor | syntax:bash, must_include(3), must_not_include(1) | 1-5 |
| 18_refactor_preserve_signature | instruction_following | syntax:python, must_include(1) | 1-5 |
| 19_git_branch_commit_push | instruction_following | syntax:bash, must_include(3), must_include_any(2) | 1-5 |
| 20_py_argparse_cli | python_automation | syntax:python, must_include(4) | 1-5 |
| 21_multifile_config_not_loaded | multi_file_reasoning | syntax:python, must_include(2), must_include_any(3) | 1-5 |
| 22_py_idempotent_insert_guard | python_automation | syntax:python, must_include(2), must_include_any(4) | 1-5 |
Coding v2 - hard set (20 cases, added 2026-07-05)
| Model | Score | Runs (range) | Pass | Latency | Temp | Measured |
| Claude Opus-4.8 (cloud reference) BEST | 0.974 | n=1 | 20/20 | 7.7s | default | 2026-07-06 |
| Ornith-1.0-397B (FP8) | 0.971 | n=3 (0.966-0.977) | 20/20 | 13.0s | default | 2026-07-05 |
| GLM-5.2 FP8 (Z.ai serving, reference) | 0.952 | n=3 (0.927-0.965) | 19/20 | 78.5s | default | 2026-07-06 |
| GLM-5.2 (NVFP4) | 0.942 | n=3 (0.918-0.968) | 20/20 | 147.6s | default | 2026-07-06 |
| Qwen3-Coder-Next-80B | 0.939 | n=3 (0.930-0.950) | 20/20 | 3.0s | default | 2026-07-05 |
| GPT-5.4 (cloud reference) | 0.935 | n=1 | 20/20 | 4.0s | default | 2026-07-06 |
| MiniMax-M3 (MXFP8) | 0.931 | n=3 (0.923-0.936) | 18/20 | 30.0s | default | 2026-07-06 |
| GPT-5.2-Codex (cloud reference) | 0.928 | n=1 | 19/20 | 7.1s | default | 2026-07-06 |
| MiniMax-M2.7 (NVFP4) | 0.901 | n=1 | 19/20 | 48.8s | default | 2026-07-05 |
| Qwen3.6-35B-A3B | 0.892 | n=3 (0.879-0.903) | 18/20 | 27.4s | default | 2026-07-06 |
The original 22-case suite saturated (the top ten sit within 0.02, below the noise
floor), so v2 was built to discriminate at the top: multi-file bug tracing (import cycles, config
precedence, systemd environment, cache-header rollouts), hard debugging (mutable defaults, DST
arithmetic, subprocess deadlock, asyncio fan-out), strict spec-compliance tasks where any missed
constraint costs points, gaps-and-islands SQL, and infrastructure tasks (Quadlet units, hotfix git
surgery, parallel-safe bash). Same scoring blend and GPT-4.1 judge as v1. The hard set separates
models the saturated suite could not: leaders that tie at 0.97-0.99 on v1 spread across 0.89-0.97
here. v1 remains the longitudinal reference; v2 decides picks.
Per-case breakdown: coding_tasks_v2.jsonl (20 cases; full prompt texts withheld to keep the benchmark uncontaminated)
| Case id | Category | Automatic checks | Judge scale |
|---|
| v2_01_interval_merge_edge | algorithms | syntax:python, must_include(3), must_include_any(2) | 1-5 |
| v2_02_toposort_cycle | algorithms | syntax:python, must_include(4), must_include_any(3) | 1-5 |
| v2_03_lru_ttl | algorithms | syntax:python, must_include(6), must_not_include(2) | 1-5 |
| v2_04_debug_mutable_default | hard_debug | syntax:python, must_include(2), must_include_any(4), must_not_include(1) | 1-5 |
| v2_05_debug_tz_dst | hard_debug | syntax:python, must_include(2), must_include_any(4), must_not_include(1) | 1-5 |
| v2_06_debug_subprocess_deadlock | hard_debug | syntax:python, must_include(4), must_not_include(1) | 1-5 |
| v2_07_debug_async_gather | hard_debug | syntax:python, must_include(2), must_include_any(2) | 1-5 |
| v2_08_spec_versioned_config | spec_compliance | syntax:python, must_include(4), must_not_include(3) | 1-5 |
| v2_09_spec_redact_logger | spec_compliance | syntax:python, must_include(6), must_include_any(2) | 1-5 |
| v2_10_spec_atomic_write | spec_compliance | syntax:python, must_include(5), must_not_include(2) | 1-5 |
| v2_11_spec_cli_exitcodes | spec_compliance | syntax:python, must_include(4), must_include_any(2), must_not_include(2) | 1-5 |
| v2_12_multifile_import_cycle | multi_file_reasoning | syntax:python, must_include(3), must_include_any(3) | 1-5 |
| v2_13_multifile_env_precedence | multi_file_reasoning | syntax:python, must_include(2), must_include_any(4) | 1-5 |
| v2_14_multifile_systemd_env | multi_file_reasoning | must_include(3), must_include_any(4) | 1-5 |
| v2_15_multifile_nginx_cache | multi_file_reasoning | must_include(3), must_include_any(4) | 1-5 |
| v2_16_sql_window_dedup | sql_hard | must_include(3), must_include_any(3) | 1-5 |
| v2_17_sql_upsert_counter | sql_hard | must_include(5), must_include_any(2) | 1-5 |
| v2_18_bash_parallel_safe | infra_hard | syntax:bash, must_include(4), must_include_any(4) | 1-5 |
| v2_19_podman_quadlet | infra_hard | must_include(7), must_include_any(2) | 1-5 |
| v2_20_gitops_hotfix | infra_hard | syntax:bash, must_include(6), must_include_any(3) | 1-5 |
Research & reasoning (27-case, none-context, task-appropriate temp=0.0-0.3)
| Model | Score | Runs (range) | Pass | Latency | Temp | Measured |
| Qwen3.6-27B (dense) BEST | 0.963 | n=1 | 26/27 | 37.6s | default | 2026-06-30 |
| Qwen3-Coder-Next-80B | 0.963 | n=1 | 26/27 | 2.0s | default | 2026-07-05 |
| Gemma-4-31B (BF16) | 0.963 | n=1 | 26/27 | 11.7s | default | 2026-06-09 |
| Gemma-4-26B-A4B (BF16) | 0.963 | n=1 | 26/27 | 2.2s | default | 2026-06-09 |
| Nex-N2-Pro (NVFP4) | 0.963 | n=1 | 26/27 | 3.2s | default | 2026-06-10 |
| DeepSeek-V4-Flash-DSpark (2x GB10) | 0.963 | n=1 | 26/27 | 25.9s | default | 2026-06-30 |
| Ornith-1.0-397B (FP8) | 0.963 | n=1 | 26/27 | 22.9s | default | 2026-07-05 |
| MiniMax-M3 (MXFP8) | 0.963 | n=1 | 26/27 | 15.6s | default | 2026-07-05 |
| gemma-4-12b-nvfp4 | 0.963 | n=1 | 26/27 | 5.0s | default | 2026-07-15 |
| gemma-4-26b-a4b-nvfp4 | 0.963 | n=1 | 26/27 | 3.5s | default | 2026-07-15 |
| DeepSeek-V4-Flash | 0.926 | n=1 | 25/27 | 15.1s | default | 2026-05-27 |
| Claude Opus-4.7 | 0.926 | n=1 | 25/27 | 9.5s | default | 2026-05-25 |
| Qwen3.6-27B BF16 (dense) | 0.926 | n=1 | 25/27 | 161.2s | default | 2026-06-09 |
| Gemma-4-12B (BF16) | 0.926 | n=1 | 25/27 | 5.6s | default | 2026-06-09 |
| Gemma-4-e4b (BF16) | 0.926 | n=1 | 25/27 | 2.7s | default | 2026-06-09 |
| GLM-5.2 (NVFP4) | 0.926 | n=1 | 25/27 | 46.1s | default | 2026-07-05 |
| GLM-5.2 FP8 (Z.ai serving, reference) | 0.926 | n=1 | 25/27 | 35.0s | default | 2026-07-06 |
| Qwen3.8-27B (FP8) | 0.926 | n=1 | 25/27 | 719.9s | default | 2026-08-14 |
| Granite-4.1-8B | 0.889 | n=1 | 24/27 | 2.6s | default | 2026-06-09 |
| Qwen3.6-35B-A3B (BF16) | 0.889 | n=1 | 24/27 | 24.2s | default | 2026-06-08 |
| devstral-small-2-24b | 0.889 | n=1 | 24/27 | 3.1s | default | 2026-07-03 |
| Gemma-4-e2b | 0.852 | n=1 | 23/27 | 12.2s | default | 2026-05-26 |
| Gemma-4-31B | 0.852 | n=1 | 23/27 | 43.2s | default | 2026-05-27 |
| Qwen3.5-Opus-distill (27B) | 0.852 | n=1 | 23/27 | 270.5s | default | 2026-05-27 |
| Ministral-3-8B | 0.852 | n=1 | 23/27 | 11.5s | default | 2026-05-27 |
| Ornith-1.0-35B (FP8) | 0.852 | n=1 | 23/27 | 8.8s | default | 2026-07-02 |
| Ornith-1.0-397B (W4A16 int4) | 0.852 | n=1 | 23/27 | 26.5s | default | 2026-07-05 |
| MiniMax-M2.7 (NVFP4) | 0.852 | n=1 | 23/27 | 12.2s | default | 2026-07-05 |
| Qwen3.6-35B-A3B | 0.815 | n=1 | 22/27 | 1.7s | default | 2026-06-02 |
| GLM-4.5-Air | 0.815 | n=1 | 22/27 | 121.8s | default | 2026-05-28 |
| Qwen3-Coder-30B | 0.815 | n=1 | 22/27 | 1.7s | default | 2026-06-09 |
| Muse-Glimmer-30B (Meta) | 0.815 | n=1 | 22/27 | 33.8s | default | 2026-08-15 |
| Gemma-4-e4b | 0.778 | n=1 | 21/27 | 20.0s | default | 2026-05-26 |
| North-Mini-Code-1.0 (FP8, Cohere) | 0.778 | n=1 | 21/27 | 6.9s | default | 2026-06-20 |
| nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.778 | n=1 | 21/27 | 4.8s | default | 2026-08-11 |
| Mistral-Small-4-119B | 0.741 | n=1 | 20/27 | 11.0s | default | 2026-05-26 |
| Mistral-Small-24B | 0.741 | n=1 | 20/27 | 4.2s | default | 2026-06-09 |
| Gemma-4-26B-A4B | 0.556 | n=1 | 15/27 | 46.4s | default | 2026-05-26 |
| Heretic-9B | 0.000 | n=1 | 0/27 | - | default | 2026-05-27 |
Re-baselined on a single cloud judge (Haiku-4.5 or GPT-4.1) per Session-13 lock-in.
Qwen3.6-35B-A3B, DeepSeek-V4-Flash and Claude Opus-4.7 tie at
0.926 - fleet reasoning is at
Opus parity. The Claude-Opus reasoning-distill (Qwen3.5-Opus, 0.852) did NOT beat native dense.
Per-case breakdown: research_tasks.jsonl (27 cases; full prompt texts withheld to keep the benchmark uncontaminated)
| Case id | Category | Automatic checks | Judge scale |
|---|
| summ_01 | long_doc_summarization | must_include(4), must_not_include(3) | 1-5 |
| summ_02 | long_doc_summarization | must_include(3), must_not_include(2) | 1-5 |
| summ_03 | long_doc_summarization | must_include(4), must_not_include(2) | 1-5 |
| summ_04 | long_doc_summarization | must_include(3), must_not_include(2) | 1-5 |
| summ_05 | long_doc_summarization | must_include(4), must_not_include(2) | 1-5 |
| summ_06 | long_doc_summarization | must_include(3), must_not_include(2) | 1-5 |
| summ_07 | long_doc_summarization | must_include(7), must_not_include(2) | 1-5 |
| mhqa_01 | multi_hop_qa | must_include(3), must_include_any(3), must_not_include(2) | 1-5 |
| mhqa_02 | multi_hop_qa | must_include(2), must_not_include(4) | 1-5 |
| mhqa_03 | multi_hop_qa | must_include(2), must_include_any(5), must_not_include(2) | 1-5 |
| mhqa_04 | multi_hop_qa | must_include(3), must_include_any(6) | 1-5 |
| mhqa_05 | multi_hop_qa | must_include(2), must_include_any(3), must_not_include(1) | 1-5 |
| mhqa_06 | multi_hop_qa | must_include_any(7), must_not_include(2) | 1-5 |
| mhqa_07 | multi_hop_qa | must_include(2), must_include_any(4) | 1-5 |
| cite_01 | citation_accuracy | must_include(3), must_include_any(2), must_not_include(4) | 1-5 |
| cite_02 | citation_accuracy | must_include_any(6), must_not_include(4) | 1-5 |
| cite_03 | citation_accuracy | must_include(1), must_include_any(2), must_not_include(3) | 1-5 |
| cite_04 | citation_accuracy | must_include_any(7), must_not_include(4) | 1-5 |
| cite_05 | citation_accuracy | must_include(4), must_include_any(2), must_not_include(3) | 1-5 |
| cite_06 | citation_accuracy | must_include_any(9), must_not_include(4) | 1-5 |
| code_01 | code_generation | syntax:bash, must_include(5), must_include_any(2), must_not_include(1) | 1-5 |
| code_02 | code_generation | syntax:python, must_include(6), must_include_any(2), must_include_all(3) | 1-5 |
| code_03 | code_generation | syntax:fish, must_include(5), must_include_any(2), must_not_include(2) | 1-5 |
| code_04 | code_generation | syntax:python, must_include(10), must_include_any(3), must_not_include(4) | 1-5 |
| code_05 | code_generation | syntax:bash, must_include(5), must_include_any(2), must_not_include(2) | 1-5 |
| code_06 | code_generation | syntax:python, must_include(7), must_include_any(2), must_not_include(1) | 1-5 |
| code_07 | code_generation | syntax:bash, must_include(10), must_include_any(2), must_not_include(1) | 1-5 |
Blog writing (10-criteria, combined = 0.5 structural + 0.5 judge, best-per-model temp)
| Model | Score | Runs (range) | Words | Judge | Gen speed | Temp | Measured |
| DeepSeek-V4-Flash BEST | 1.000 | n=2 | 3774 | - | 64s | 0.3 | 2026-05-28 |
| Qwen3.6-27B BF16 (dense) | 1.000 | n=1 | 3829 | 10/10 | 477s 28 t/s | default | 2026-06-08 |
| Nex-N2-Pro (NVFP4) | 1.000 | n=1 | 4411 | 10/10 | 95s 91 t/s | default | 2026-06-10 |
| Qwen3.6-27B (dense) | 1.000 | n=1 | 3305 | 10/10 | 81s 98 t/s | default | 2026-06-30 |
| qwen3.6-27b-bf16 | 1.000 | n=3 (1.000-1.000) | 3157 | 10/10 | 148s 60 t/s | default | 2026-07-02 |
| Ornith-1.0-397B (FP8) | 1.000 | n=3 (1.000-1.000) | 3857 | 10/10 | 64s 120 t/s | default | 2026-07-05 |
| GLM-5.2 (NVFP4) | 1.000 | n=1 | 3439 | 10/10 | 207s 40 t/s | default | 2026-07-06 |
| MiniMax-M3 (MXFP8) | 1.000 | n=3 (1.000-1.000) | 3317 | 10/10 | 46s 123 t/s | default | 2026-07-05 |
| Claude Opus-4.8 (cloud reference) | 1.000 | n=1 | 3297 | 10/10 | 125s 77 t/s | default | 2026-07-06 |
| Claude Fable 5 (cloud reference) | 1.000 | n=1 | 3630 | 10/10 | 135s 81 t/s | default | 2026-07-06 |
| Ornith-1.0-35B (FP8) | 0.963 | n=3 (0.945-1.000) | 2210 | 10/10 | 66s 166 t/s | default | 2026-07-02 |
| gpt-oss-120b | 0.945 | n=2 | 1938 | - | 132s | 0.0 | 2026-05-28 |
| Gemma-4-31B (BF16) | 0.945 | n=1 | 2390 | 10/10 | 169s 24 t/s | default | 2026-06-09 |
| Gemma-4-12B (BF16) | 0.945 | n=1 | 2291 | 10/10 | 78s 53 t/s | default | 2026-06-09 |
| North-Mini-Code-1.0 (FP8, Cohere) | 0.945 | n=1 | 2269 | 10/10 | 43s 139 t/s | default | 2026-06-20 |
| DeepSeek-V4-Flash-DSpark (2x GB10) | 0.945 | n=1 | 2646 | 10/10 | 131s 46 t/s | default | 2026-06-30 |
| Laguna-S-2.1-NVFP4 | 0.945 | n=1 | 3255 | 10/10 | 55s 106 t/s | default | 2026-07-26 |
| Ornith-1.0-397B (W4A16 int4) | 0.945 | n=3 (0.944-0.945) | 3251 | 9/10 | 70s 104 t/s | default | 2026-07-03 |
| GLM-4.7-Flash | 0.945 | n=2 | 2640 | - | 59s | 0.0 | 2026-05-28 |
| Qwen3.6-35B-A3B (BF16) | 0.944 | n=1 | 3293 | 9/10 | 49s 171 t/s | default | 2026-06-08 |
| GLM-5.2 (IQ2_M 2-bit) | 0.944 | n=1 | 4127 | 9/10 | 199s 39 t/s | default | 2026-06-20 |
| GPT-5.4 (cloud reference) | 0.944 | n=1 | 3246 | 9/10 | 64s 102 t/s | default | 2026-07-06 |
| MiniMax-M2.7 (NVFP4) | 0.907 | n=3 (0.833-1.000) | 3966 | 9/10 | 50s 129 t/s | default | 2026-07-05 |
| Qwen3.6-35B-A3B (FP8) | 0.889 | n=2 | 2597 | - | 38s | 0.0 | 2026-05-28 |
| Gemma-4-26B-A4B | 0.889 | n=1 | 2280 | 9/10 | 128s 44 t/s | default | 2026-05-26 |
| Gemma-4-31B | 0.889 | n=1 | 2063 | 9/10 | 553s 7 t/s | default | 2026-05-27 |
| Gemma-4-26B-A4B (BF16) | 0.889 | n=1 | 2128 | 10/10 | 26s 147 t/s | default | 2026-06-09 |
| Gemma-4-e4b (BF16) | 0.889 | n=1 | 2676 | 10/10 | 40s 122 t/s | default | 2026-06-09 |
| Laguna-XS-2.1-NVFP4 | 0.889 | n=1 | 1703 | 9/10 | 30s 192 t/s | default | 2026-07-26 |
| Qwen3-Coder-Next-80B | 0.870 | n=3 (0.833-0.945) | 3210 | 9/10 | 36s 180 t/s | default | 2026-07-05 |
| gemma-4-26b-a4b-nvfp4 | 0.870 | n=3 (0.833-0.945) | 1827 | 9/10 | 34s 97 t/s | default | 2026-07-15 |
| Qwen3.6-27B-NVFP4 (unsloth, dense) | 0.855 | n=5 (0.333-1.000) | 3211 | 9/10 | 89s 98 t/s | default | 2026-07-12 |
| Gemma-4-e2b | 0.833 | n=1 | 2864 | 9/10 | 80s 78 t/s | default | 2026-05-26 |
| command-a-plus | 0.833 | n=1 | 1953 | 9/10 | 103s 49 t/s | default | 2026-05-26 |
| Qwen3.5-Opus-distill (27B) | 0.833 | n=1 | 4197 | 9/10 | 240s 37 t/s | default | 2026-05-26 |
| Laguna-M.1-NVFP4 | 0.833 | n=1 | 1926 | 9/10 | 58s 79 t/s | default | 2026-07-26 |
| nemotron-3.5-lightning-30b-a3b-nvfp4 | 0.833 | n=1 | 3582 | 9/10 | 104s 73 t/s | default | 2026-08-11 |
| Claude Opus-4.7 | 0.805 | n=2 | 2342 | - | 107s | default (~0.0) | 2026-05-28 |
| Qwen3.6-35B-A3B-NVFP4-Fast (unsloth) | 0.800 | n=5 (0.278-0.945) | 2719 | 10/10 | 56s 138 t/s | default | 2026-07-12 |
| devstral-small-2-24b | 0.796 | n=3 (0.778-0.833) | 1307 | 9/10 | 32s 124 t/s | default | 2026-07-03 |
| gpt-oss-20b | 0.778 | n=1 | 1793 | 9/10 | 34s 243 t/s | default | 2026-05-25 |
| GLM-4-9B | 0.778 | n=1 | 898 | 7/10 | 12s 181 t/s | default | 2026-05-25 |
| Mixtral-8x22B | 0.778 | n=1 | 931 | 7/10 | 41s 60 t/s | default | 2026-05-26 |
| Granite-4.1-8B | 0.778 | n=1 | 868 | 9/10 | 13s 169 t/s | default | 2026-06-09 |
| Qwen3.8-27B (FP8) | 0.778 | n=1 | 2522 | 8/10 | 700s 8 t/s | default | 2026-08-14 |
| Mistral-Small-24B | 0.722 | n=1 | 1342 | 7/10 | 39s 74 t/s | default | 2026-06-09 |
| Gemma-4-e4b | 0.722 | n=1 | 2434 | 7/10 | 105s 47 t/s | default | 2026-05-26 |
| Mistral-Small-4-119B | 0.722 | n=1 | 3262 | 8/10 | 281s 28 t/s | default | 2026-05-26 |
| gemma-4-12b-nvfp4 | 0.704 | n=3 (0.333-0.889) | 6138 | 3/10 | 202s 59 t/s | default | 2026-07-15 |
| Mistral-Medium-3.5-128B | 0.611 | n=1 | 1469 | 7/10 | 450s 18 t/s | default | 2026-05-26 |
| GLM-4.5-Air | 0.500 | n=1 | 4541 | 3/10 | 1360s 6 t/s | default | 2026-05-25 |
| nemotron-3-nano:30b | 0.445 | n=1 | 5618 | 2/10 | 156s 51 t/s | default | 2026-05-25 |
| Qwen3.6-35B-A3B | 0.278 | n=1 | 195 | 3/10 | 212s 38 t/s | default | 2026-05-25 |
| Mixtral-8x7B | 0.222 | n=1 | 79 | 2/10 | 1s 137 t/s | default | 2026-05-26 |
| Qwen3-Coder-30B | 0.111 | n=1 | 0 | 1/10 | 11s 0 t/s | default | 2026-06-09 |
3,000+ word SEO CMMC blog from a fixed prompt. Each row reports the model's
best-scoring temperature when the Session-14 Bench-2 sweep covered it (the 2026-05-28 model
set only); models benched after that date show their single measured temperature. Scores within
0.03 are statistical ties. Autoblog runs overnight in batch, so quality decides.
Blog temperature sweep (Session-14 Bench-2, judge=GPT-4.1, N=2 reruns per cell)
| Model | Temp | Mean score | Range (min-max) | Mean words | Mean gen |
| Claude Opus-4.7 | default (~0.0) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| Claude Opus-4.7 | default (~0.3) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| Claude Opus-4.7 | default (~0.7) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| Claude Opus-4.7 | default (~1.0) | 0.805 | 0.778-0.833 | 2342 | 107.2s |
| DeepSeek-V4-Flash | 0.0 | 0.972 | 0.944-1.000 | 4194 | 78.4s |
| DeepSeek-V4-Flash | 0.3 BEST | 1.000 | 1.000-1.000 | 3774 | 64.2s |
| DeepSeek-V4-Flash | 0.7 | 1.000 | 1.000-1.000 | 3732 | 65.0s |
| DeepSeek-V4-Flash | 1.0 | 0.972 | 0.944-1.000 | 4025 | 72.4s |
| GLM-4.7-Flash | 0.0 BEST | 0.945 | 0.944-0.945 | 2640 | 58.8s |
| GLM-4.7-Flash | 0.3 | 0.805 | 0.722-0.889 | 2210 | 53.1s |
| GLM-4.7-Flash | 0.7 | 0.861 | 0.833-0.889 | 2748 | 57.9s |
| GLM-4.7-Flash | 1.0 | 0.889 | 0.889-0.889 | 2072 | 49.5s |
| Qwen3.6-35B-A3B (FP8) | 0.0 BEST | 0.889 | 0.889-0.889 | 2597 | 37.5s |
| Qwen3.6-35B-A3B (FP8) | 0.3 | 0.806 | 0.667-0.945 | 3204 | 39.8s |
| Qwen3.6-35B-A3B (FP8) | 0.7 | 0.584 | 0.445-0.722 | 2517 | 44.2s |
| Qwen3.6-35B-A3B (FP8) | 1.0 | 0.611 | 0.611-0.611 | 2549 | 41.5s |
| gpt-oss-120b | 0.0 BEST | 0.945 | 0.945-0.945 | 1938 | 131.8s |
| gpt-oss-120b | 0.3 | 0.889 | 0.889-0.889 | 2145 | 137.4s |
| gpt-oss-120b | 0.7 | 0.945 | 0.945-0.945 | 2144 | 127.8s |
| gpt-oss-120b | 1.0 | 0.917 | 0.889-0.945 | 1710 | 124.2s |
Session-14 Bench-2 temperature sweep measured 2026-05-28 13:48 EDT. Reasoning models (DeepSeek-V4-Flash, Qwen3.6-35B-A3B) may treat temperature
as a hint during their reasoning phase; a flat row across temps is itself a finding. Claude Opus-4.7
rejects the temperature parameter; its rows are copied from a single default-temperature run.
How the scores are produced
Coding - 22 real operations tasks, not textbook puzzles. Drawn from Petronella
Technology Group's day-to-day infrastructure and SEO work: 8 Python automation tasks (OpenAI-compatible
API clients, retry/backoff decorators, concurrent endpoint health checks, log parsing, argparse CLIs),
3 bash/fish shell-ops tasks (timestamped backups, safe rsync deploys, fleet uptime sweeps), 3 production
config edits (systemd OnCalendar timers, YAML inventories, .htaccess 301 rules), 3 instruction-following
traps (tasks that fail if a stated constraint is ignored, such as preserving a function signature),
2 debug-and-fix cases, 1 SQL analytics query, 1 shell refactor, and 1 multi-file reasoning case.
Each case is scored 0.6 x automatic checks (required content, regex gates, and real syntax validation:
bash -n, python compile, fish -n) + 0.4 x a blind LM-judge
correctness score (1-5; the judge sees only task and solution, never the model's identity).
Deterministic decoding (temp 0.0). The leaderboard number is the mean across all 22 cases.
Research & reasoning - 27 cases in four categories. 7 long-document
summarization cases (word bounds, bullet counts, must-include / must-not-include gates against supplied
source documents), 7 multi-hop QA cases (the answer requires chaining facts across documents),
6 citation-grounding cases (presence and format of citations regex-verified against the supplied sources; semantic citation correctness is measured separately in PTG's cited-RAG evaluations, which score a much harsher 0.65-0.80 F1), and 7 grounded
code-generation cases with syntax gates. Same 0.6 auto + 0.4 blind-judge blend as coding.
Blog writing - one fixed brief, structural gates plus publish-readiness. Every model
gets the identical brief: a 3,000+ word CMMC compliance post in clean HTML. Nine automatic structural
gates (word count, H1/H2/H3 hierarchy with 8+ sections, no leftover markdown artifacts, FAQ with 5+
substantive Q&As, a well-formed comparison table, internal links, correct call-to-action, and brand
rules) plus a 10-criteria publish-readiness judge score (1-10) covering factual accuracy (CMMC 2.0
levels, NIST 800-171), tone, and SEO structure. Combined = 0.5 x structural pass rate + 0.5 x judge.
Blog rows report each model's best temperature from the Session-14 sweep where covered.
Voice tool-calling - 88-case corpus with decline-safety distractors. Real assistant
tool schemas; the corpus mixes valid calls with hard distractors that look like tool calls but must be
refused. Scored on correct calls, correct declines, and hallucinated-call rate. This suite decides the
voice tool-caller pick; decline-safety on distractors is the deciding metric.
Agentic coding - SWE-bench Verified, no LM judge. Stratified 150-instance subset of
SWE-bench Verified run through the official harness with mini-swe-agent; a task counts only if the
generated patch resolves the issue's test suite. Referenced in picks and findings (for example
DeepSeek-V4-Flash 67.3%, measured 2026-06-28); not a leaderboard column.
Throughput - measured on the serving hardware. Single-stream (c=1) plus concurrent
c=8 and c=32 sweeps of 800-token generations against the exact serve config named in each finding.
Speeds quoted in findings are decode tokens per second on the stated GPUs and quant.
Judging integrity - the cross-judge gate. A generator never grades itself:
self-judging inflated scores 20-37% in our measurements. The campaign judge is locked to GPT-4.1
(code-enforced since 2026-05-30, with a guard that flags any same-family judge pairing). Legacy rows
graded by Claude Haiku-4.5 are labeled; DeepSeek-V4-Flash is accepted for bulk grading after validating
equivalence with GPT-4.1 (r=0.962, mean absolute delta 0.18 on the 1-5 scale; re-audited 2026-07-06 on 40 blog outputs: mean abs delta 0.45 on the 1-10 scale, 87.5% within 1 point, second judge slightly stricter). Excluded from all
tables: misconfigured or errored runs, runs without a cross-family judge, and hardware-specific Mac
coding runs (see the Apple Silicon matrix).
Reproducibility. Every row carries an absolute measurement date; the exact
quantization is recorded in the run label (quant changes the result: the same model at FP8, NVFP4, and
int4 scores differently and is listed separately). Headline claims run N=3 to N=5 repetitions. As of 2026-07-05 the board shows the MEAN of each
model's largest rep family (newest on ties) with n and min-max range in the Runs column (previously best-run, which
rewarded models benched more often); scores within 0.03 are statistical ties.
Findings
Measured 2026-07-15
Is Gemma-4-26B-A4B better than
Qwen3.6, and can it run in the agent gateway? It depends on task difficulty, judged by a
three-model frontier panel (GPT-5.4 + Claude Opus-4.8 + GLM-5.2, median score on identical
generations, N=5). On easy/routine coding Qwen wins (dense 27B 0.983, 35B-FP8 0.979 vs Gemma
0.962); on the hard coding set Gemma wins decisively (0.915 vs every Qwen variant 0.78-0.81),
corroborated by the deterministic auto-checks. Research and blog are ties (about 0.86 and 0.79).
Agentic: served with the correct vLLM tool parser (gemma4, not pythonic - the wrong one produces
zero tool calls), Gemma is the strongest tool-caller (55 correct calls vs Qwen's 51) and solves
every multi-turn task (pass@max 1.00 vs Qwen-FP8 0.778). Both models were verified end-to-end
through the real agent gateway - writing and executing code and returning correct results - which
requires serving the model at 131k context (the agent system prompt is about 58k tokens). NVFP4 is
near-lossless for Gemma-4 but not for Qwen3.6-35B (a long-form runaway appears at FP4; use FP8).
Measured 2026-07-06
Does GLM-5.2 lose anything at FP4,
and is the six-GPU FP8 layout justified? No measurable loss. GLM-5.2 at FULL FP8 fidelity
(Z.ai's own serving, via OpenRouter) scored 0.952 on the v2 hard set (n=3: 0.927/0.964/0.965) and
0.92 on research - versus the FP4 quant on the fleet's 4-way bridge at 0.942 and 0.94. The FP8
delta (+0.010 coding, -0.02 research) is inside the 0.03 tie threshold: FP4 is effectively
lossless for GLM-5.2 on these tests, and both configurations score below Ornith-1.0-397B-FP8
(0.971). Implication for fleet layout: an FP8 deployment spanning six GPUs and both NVLink
bridges buys no measurable quality over FP4 on four GPUs.
Measured 2026-07-06
Do two independent judges agree on
these scores? Yes. A dual-judge audit re-graded 40 stored blog outputs with DeepSeek-V4-Flash
(served on the DGX Spark pair) against their original GPT-4.1 scores using the identical rubric:
mean absolute difference 0.45 points on the 1-10 scale, 87.5% of pairs within one point, and the
second judge ran slightly STRICTER on average (signed mean -0.3) with its five largest
disagreements spread across Anthropic, MiniMax, and Qwen outputs - no family favoritism pattern.
This is the periodic cross-judge integrity check the methodology commits to.
Measured 2026-07-06
How does the fleet compare to the
July 2026 cloud frontier on the same tests? Cloud reference rows added: Claude Opus-4.8 v1
0.995 / v2 0.974 / blog 1.000 (GPT-4.1 judge); GPT-5.4 v1 0.970 / v2 0.935 / blog 0.944 and
GPT-5.2-Codex v1 0.982 / v2 0.928 (both Haiku-judged per the cross-family gate; OpenAI rows and
Anthropic/fleet rows use different judges, so treat cross-vendor deltas under 0.05 with care).
Read: fleet-owned Ornith-1.0-397B-FP8 (v2 0.971, zero marginal cost, CUI-capable) sits between
GPT-5.4 (0.935) and Opus-4.8 (0.974) on the hard set. Also measured: Claude Fable 5's API safety
layer CONTENT-FILTERED 9 of 22 routine infrastructure tasks (finish_reason=content_filter after
~3 tokens: rsync deploy scripts, a phone-number regex, an argparse CLI), making it unusable as a
coding baseline here; its blog run scored 1.000. GPT-5.2-Codex is Responses-API-only (the harness
gained --use-responses-api).
Measured 2026-07-06
Do the bridge picks survive a harder
test set and multi-topic writing? Yes, and the ordering sharpens. The 20-case v2 coding set
(built because the original suite saturated) spreads the leaders decisively: Ornith-1.0-397B-FP8
0.971 (n=3, 20/20 every rep, 13 s per case), GLM-5.2-NVFP4 0.942 (n=3, high variance 0.918-0.968
at 150 s per case), Qwen3-Coder-Next 0.939 (n=3, 2.6 s per case), MiniMax-M3 0.931 (n=3, drops 1-2
cases per rep), incumbent Qwen3.6-35B-A3B 0.892 (n=3). MiniMax-M3's v1 crown (0.984) inverts on
v2: the saturated suite was measuring judge quibbles, not capability. Blog re-tested at 3 topics x
3 reps: Ornith-397B a perfect 1.000 on all nine runs; MiniMax-M3 0.975 mean (short only on the
CMMC brief); MiniMax-M2.7 0.94 mean (its earlier single-run 1.000 was variance); Qwen3-Coder-Next
0.887 mean. Qwen3-Coder-Next also posted 85.2% on the 88-case tool corpus (53 correct calls, 24
correct declines, 1 hallucinated call): strong for agentic coding, not voice-grade (gpt-oss-20b
remains 0% hallucination). Remaining before any production routing change: a local SWE-bench
Verified run for Qwen3-Coder-Next.
Measured 2026-07-05
Which open model is best on each NVLink
island of the 6x H200 fleet? First head-to-head since the 4-way bridge install. On the
4-way bridge (575 GB), Ornith-1.0-397B-FP8 (MIT) is the balanced winner: coding 0.980
(N=3), blog 1.000, research 0.91 at 119.7 t/s single-stream and 639 t/s at c=8, roughly 2x the speed
of speed-tuned GLM-5.2-NVFP4 at equal or better quality on two of three axes. MiniMax-M3-MXFP8
posted the campaign's best coding rep (0.991; N=3 mean 0.984) with 1M context and multimodal input,
held back only by its non-MIT community license. GLM-5.2-NVFP4 keeps the research crown (0.94,
citation accuracy 0.992). On the 2-way bridge (287 GB), Qwen3-Coder-Next-FP8 (80B-A3B)
is the coder/agentic pick: 0.977 coding with zero variance across three full reps, clean native tool
calls, 178 t/s single-stream and 2,789 t/s aggregate at c=32. Ruled out: Kimi-K2.6 (594 GB INT4
exceeds any island), GLM-5.x-FP8 (754 GB needs all six cards), Nemotron-3-Ultra-NVFP4 (numerically
broken on Hopper: loads but emits gibberish). Serving notes: Ornith-397B-FP8 requires
VLLM_TEST_FORCE_FP8_MARLIN=1 plus an explicit chat template; MiniMax-M3 requires
--block-size 128; the 2026-07-03 vLLM nightly has broken block-scaled FP8 kernels on
Hopper (pin v0.24.0).
Which models do we keep resident and route to? Coding → Qwen3.6-35B-A3B (0.975, GPT-4.1).
Research → Qwen3.6-35B (speed) / DeepSeek-V4-Flash (long-context), both Opus-parity. Blog →
DeepSeek-V4-Flash / GLM-4.7-Flash. Voice tool-call → gpt-oss-20b. Edge appliance → Granite-4.1-8B
(+ Gemma-4-e4b for tool-driving). Reserve Claude Opus for the top few percent.
Do self-improving loops help small models? A loop is a capability amplifier, not an
equalizer: Qwen3.6-35B goes 0.917→1.0 with iterations; small models (Gemma-4-e4b, GLM-4.7-Flash)
stay flat at 0.917. The done-gate makes a small model honest (no silent fabrication), not capable.
Self-improving loop vs a general agent (pi.dev) on a 4B? Gemma-4-e4b scored 0.917 in a
done-gated loop vs 0.0 in pi.dev (it fabricated all 12 tasks). Air-gap appliances should pair a small
model with a programmatic verifier loop, never a general agent framework.
Are there other open models worth adding? A live Feb–May 2026 scan found none that beat
the incumbents; Qwen3.7/Qwen4, DeepSeek-R2, Phi-5 and Grok-3 weights are unreleased or hosted-only.
What is the best model for a single RTX PRO 6000 96GB (Blackwell) card, and is a 35B a waste of it?
No displacer. Qwen3.6-35B-A3B (3B active) is the best all-rounder that fits one card: it wins research and
cited-RAG outright and leads coding (0.975, GPT-4.1). The models large enough to “fill” the card
(gpt-oss-120b 0.964, Mistral-Small-4-119B 0.957) are slower and weaker on the role’s core axes. A
low-active MoE is the correct shape for a 96GB concurrency server: comparable NVFP4 models scale to ~2,000
t/s aggregate at c=32 on this card. Spare VRAM is best spent on KV/concurrency, or on NVFP4 (same quality at half
the VRAM, freeing room to co-locate a second model), not on a bigger-but-worse model. Mistral-Small-4-119B is the
lone alternative, and only if the card is redefined as a cited-RAG / vision / compliance resident.
Is a purpose-built Rust inference engine (Atlas) faster than our tuned vLLM on Blackwell? (measured 2026-06-07)
No. On identical GB10 (DGX-Spark-class) hardware and the same Qwen3.6-35B-A3B-NVFP4 model, our tuned vLLM
(NVIDIA MTP recipe) ran 116–119 tok/s steady-state vs Atlas’s 88.9; Atlas’s advertised
“130–133 tok/s” and “3.1× faster than vLLM” did not reproduce (the 3.1× is vs an
untuned vLLM). Atlas serving is quality-preserving - blog 0.944 (ties our blog leader) and
6/6 on a coding spot-check - and ships an ~8×-smaller (2.98 GB) no-Python single binary. That makes it a
candidate packaging vehicle for an air-gapped compliance appliance, not a throughput upgrade. Its multi-node
expert-parallel mode is not yet shipping (runtime is single-node only).
Does an agentic multi-hop retriever beat single-shot RAG for compliance Q&A? (measured 2026-06-07)
On hard multi-hop CMMC / NIST 800-171 questions, an RL-trained search agent (Harness-1, 21B, gpt-oss-20b base)
found every gold control (retrieval recall 1.000) where single-shot dense top-8 reached only 0.881 -
it recovers the deep 2nd/3rd-hop controls single-shot drops at production cutoffs. But its curated answer (0.929)
only matched single-shot top-15 (the curation step, not the search, is the bottleneck) and cost ~1,000× the
latency - so the value is exhaustive batch retrieval (audit / SSP gap analysis), not interactive RAG.
Control-id deduplication remains the cheap universal lever: it lifts both single-shot (0.786→0.881) and the
agent (0.905→1.000).
Benchmarked and published by Petronella Technology Group, Inc. All scores are
produced by the PTG llm-benchmark harness described in "How the scores are produced" above.