PTG Fleet Model Leaderboards

Open-weight models benchmarked on Petronella Technology Group's private GPU fleet (6x H200 NVL, RTX PRO 6000 Blackwell, GB10). Every published score is cross-judged: the generator never grades itself, and the campaign judge is locked to GPT-4.1 (Session-13/15 lock-in; Haiku-4.5 or validated-equivalent DeepSeek-V4-Flash on legacy rows). Last updated: 2026-08-15 12:45 EDT. Auto-generated from harness result files; the per-row Measured column is the data-collection date.
Current picksCodingCoding v2Research BlogTemp sweepFindingsMethod Throughput dashboardApple Silicon matrix

Current picks by hardware role (verdict 2026-07-05; re-verified 2026-07-06 on the v2 hard set)

4-way H200 NVL bridge (575 GB)
Ornith-1.0-397B-FP8 (MIT): leads the v2 hard set outright (0.971, n=3, 20/20 every rep) and scored a perfect blog 1.000 on all nine runs across three topics; research 0.91 at 120 t/s single-stream. Alternatives: MiniMax-M3-MXFP8 (v2 0.931; 1M ctx, multimodal; restrictive community license) and GLM-5.2-NVFP4 (v2 0.942 but 150 s per case with high variance; still the research leader at 0.94).
2-way H200 NVL bridge (287 GB)
Qwen3-Coder-Next-FP8 (80B-A3B, Apache-2.0): best coder in the 2-way class on the v2 hard set (0.939, n=3, at 2.6 s per case; the nearest class rival needs 20x the latency for less), 85.2% on the 88-case tool corpus, 178 t/s single-stream and 2,789 t/s aggregate at c=32. Long-form writer alternative: MiniMax-M2.7-NVFP4 (blog 0.94 mean across three topics, v2 0.901).
Single 96 GB card
Qwen3.6-35B-A3B-FP8: best all-rounder that fits one card (coding 0.975, wins research and cited-RAG); no displacer found (2026-05-31).
Long context (524K)
DeepSeek-V4-Flash: coding 0.964, blog 0.944, SWE-bench Verified 67.3% measured locally (2026-06-28); also the fleet's validated free grader.
Voice tool-caller
gpt-oss-20b: 125/125 correct declines, 0% tool hallucination on the 88-case corpus (2026-05-29).
Air-gap edge appliance
Ministral-3-8B (2026-05-27 re-test); Granite-4.1-8B and Gemma-4-e4b for tool-driving. All Mistral models need vLLM --tool-call-parser mistral.

Coding v1 (22 PTG ops/SEO/debug tasks; longitudinal reference - saturated at the top)

ModelScoreRuns (range)PassLatencyTempMeasured
Claude Opus-4.8 (cloud reference) BEST0.995n=122/222.6sdefault2026-07-06
Claude Opus-4.70.991n=122/222.7sdefault2026-05-25
Laguna-S-2.1-NVFP40.991n=122/223.3sdefault2026-07-27
Qwen3.6-27B BF16 (dense)0.990n=120/2266.4sdefault2026-06-08
gpt-oss:20b0.989n=122/222.5sdefault2026-05-28
DeepSeek-V4-Flash-DSpark (2x GB10)0.986n=122/2211.2sdefault2026-06-30
Qwen3.6-27B (dense)0.984n=122/2225.2sdefault2026-06-30
MiniMax-M3 (MXFP8)0.984n=3 (0.975-0.991)21/225.2sdefault2026-07-05
Laguna-M.1-NVFP40.984n=122/226.4sdefault2026-07-26
DeepSeek-V4-Pro (cloud reference)0.982n=122/224.1sdefault2026-06-28
GPT-5.2-Codex (cloud reference)0.982n=122/222.6sdefault2026-07-06
Ornith-1.0-397B (FP8)0.980n=3 (0.977-0.982)22/222.7sdefault2026-07-05
Qwen3.6-35B-A3B-NVFP4-Fast (unsloth)0.980n=5 (0.971-0.986)22/2220.0sdefault2026-07-12
Gemma-4-31B0.980n=122/2228.9sdefault2026-05-27
Nex-N2-Pro (NVFP4)0.980n=122/2218.1sdefault2026-06-10
North-Mini-Code-1.0 (FP8, Cohere)0.980n=5 (0.976-0.986)22/223.4sdefault2026-06-20
Qwen3-Coder-Next-80B0.977n=3 (0.977-0.977)22/220.6sdefault2026-07-05
Gemma-4-12B0.977n=122/2215.8sdefault2026-06-03
Qwen3.8-27B (BF16)0.977n=122/2235.7sdefault2026-08-14
MiniMax-M2.7 (NVFP4)0.977n=3 (0.965-0.984)22/2217.3sdefault2026-07-05
gemma-4-26b-a4b-nvfp40.977n=3 (0.970-0.980)22/221.5sdefault2026-07-15
Qwen3.6-35B-A3B0.975n=5 (0.961-0.984)22/220.6sdefault2026-05-29
gpt-oss-20b0.975n=122/2211.8sdefault2026-05-28
ornith-35b0.975n=14/226.6sdefault2026-06-29
glm-5.20.975n=122/226.3sdefault2026-07-13
GLM-4.7-Flash0.973n=122/228.9sdefault2026-05-28
Qwen3.6-35B-A3B (BF16)0.973n=122/2222.7sdefault2026-06-08
GLM-5.2 (NVFP4)0.972n=5 (0.959-0.986)22/2246.0sdefault2026-07-03
Ornith-1.0-397B (W4A16 int4)0.970n=5 (0.959-0.982)21/223.6sdefault2026-07-03
Gemma-4-26B-A4B (BF16)0.970n=122/221.0sdefault2026-06-09
GPT-5.4 (cloud reference)0.970n=122/221.7sdefault2026-07-06
Laguna-XS-2.1-NVFP40.970n=121/220.7sdefault2026-07-26
Mistral-Medium-3.5-128B0.968n=122/224.6sdefault2026-05-26
Qwen3.6-27B-NVFP4 (unsloth, dense)0.967n=5 (0.966-0.970)21/2214.4sdefault2026-07-12
qwen3-coder:480b0.966n=122/222.1sdefault2026-06-28
Qwen3.8-27B (FP8)0.966n=121/22124.6sdefault2026-08-14
gpt-oss-120b0.964n=121/228.4sdefault2026-05-28
DeepSeek-V4-Flash0.964n=5 (0.961-0.970)21/220.9sdefault2026-05-29
GLM-5.2 (IQ2_M 2-bit)0.959n=121/2226.9sdefault2026-06-20
nemotron-3.5-lightning-30b-a3b-nvfp40.958n=121/2218.8sdefault2026-08-11
gpt-4.10.957n=121/221.4sdefault2026-05-25
Mistral-Small-4-119B0.957n=121/221.4sdefault2026-05-27
gemma4-coder0.957n=122/223.1sdefault2026-07-14
GLM-4.5-Air0.956n=117/22105.3sdefault2026-05-28
claude-haiku-4-5-202510010.956n=5 (0.955-0.959)21/221.2sdefault2026-05-29
Gemma-4-31B (BF16)0.955n=121/224.9sdefault2026-06-09
z-ai/glm-5.20.955n=121/2217.3sdefault2026-06-27
Qwen3-Coder-30B0.952n=122/220.8sdefault2026-06-09
openai/gpt-oss-20b0.952n=121/221.8sdefault2026-05-26
qwen3-coder:30b0.952n=122/220.8sdefault2026-05-27
Gemma-4-12B (BF16)0.952n=121/222.5sdefault2026-06-09
Gemma-4-e4b (BF16)0.949n=122/221.8sdefault2026-06-09
Muse-Glimmer-30B (Meta)0.949n=121/2214.6sdefault2026-08-15
Ornith-1.0-35B (FP8)0.948n=5 (0.913-0.970)21/223.7sdefault2026-07-02
Qwen3.8-27B (NInfer int4+MTP3)0.941n=121/2218.8sdefault2026-08-15
command-a-plus0.939n=121/2218.1sdefault2026-05-26
Gemma-4-e2b0.935n=121/2212.0sdefault2026-05-26
Qwen3.5-Opus-distill (27B)0.934n=121/2213.1sdefault2026-05-26
Qwen3.8-27B (Q4_K_XL)0.934n=120/22171.1sdefault2026-08-15
Devstral-Small-2-24B0.932n=120/226.7sdefault2026-05-28
gemma-4-12b-nvfp40.931n=3 (0.930-0.932)21/222.1sdefault2026-07-15
Gemma-4-e4b0.931n=120/2213.6sdefault2026-05-26
Falcon-H1R-7B0.926n=119/2284.2sdefault2026-05-27
devstral-small-2-24b0.925n=5 (0.900-0.936)20/220.7sdefault2026-07-03
Mixtral-8x22B0.909n=121/221.9sdefault2026-05-26
Mistral-Small-24B0.909n=120/221.4sdefault2026-06-09
Granite-4.1-8B0.903n=119/220.8sdefault2026-06-09
hf.co/tiiuae/Falcon-H1R-7B-GGUF:Q4_K_M0.885n=119/2249.9sdefault2026-05-27
q25-gptq0.885n=119/222.4sdefault2026-05-29
Apriel-1.6-15B0.880n=119/22110.6sdefault2026-05-27
Gemma-4-26B-A4B0.877n=118/2218.4sdefault2026-05-26
ornith-9b0.875n=119/2254.0sdefault2026-06-29
Ministral-3-8B0.872n=118/222.9sdefault2026-05-27
hf.co/unsloth/Ministral-3-8B-Instruct-2512-GGUF:Q4_K_M0.858n=118/220.8sdefault2026-05-27
qwen2.5:7b-instruct0.847n=118/222.1sdefault2026-05-29
l31-awq0.847n=119/223.4sdefault2026-05-29
LFM2.5-8B-A1B (Liquid)0.846n=117/222.0sdefault2026-06-03
q25-awq0.820n=117/222.5sdefault2026-05-29
llama3.1:8b0.820n=117/222.7sdefault2026-05-29
l31-gptq0.794n=115/222.9sdefault2026-05-29
Mixtral-8x7B0.773n=117/220.7sdefault2026-05-26
MiniCPM5-1B0.644n=110/226.2sdefault2026-05-27
Incumbent fleet coder Qwen3.6-35B-A3B now co-leads. North-Mini-Code-1.0 (Cohere, Apache-2.0, 30B-total/3B-active MoE) matches the very top on PTG coding - 0.980 ± 0.004 (N=5), statistically tied with Nex-N2-Pro and edging Qwen3.6-35B-FP8 (0.975) - and loop-amplifies on agentic tasks (cap=1 0.833 → cap=5 0.917, fabrication 2→1), unlike the non-amplifiers below. Apache-2.0, FP8 fits one card, fast, clean tools; a genuine alternative fleet coder (Session-22). It is not a research/RAG model (0.821) and blogs short of 3,000 words. Gemma-4-12B (0.977) ties the 31B/Qwen3.6 on single-shot coding but is single-shot only - it does NOT loop-amplify (cap=1=cap=5=0.917 on agentic tasks) and fabricates "done" under loop pressure; use it as a fast coder, not an agent. LFM2.5-8B-A1B is an on-device model (edge-tier 0.846, 3.41% tool-call hallucination) - not a fleet upgrade. Score = mean of each model's largest rep family (newest on ties); scores within 0.03 of each other are statistical ties. Claude Opus-4.7 reference = 1.00. Scores cross-judged by GPT-4.1, Haiku-4.5, or validated-equivalent DeepSeek-V4-Flash. Temperature column: stamped value where present; "default" = pre-Session-13 runs (eval_coding.py default 0.0).
Per-case breakdown: coding_tasks.jsonl (22 cases; full prompt texts withheld to keep the benchmark uncontaminated)
Case idCategoryAutomatic checksJudge scale
01_bash_backup_recent_htmlbash_opssyntax:bash, must_include(3), must_include_any(3)1-5
02_fish_fleet_uptimebash_opssyntax:fish, must_include(3), must_include_any(2)1-5
03_bash_safe_rsync_deploybash_opssyntax:bash, must_include(3), must_include_any(2)1-5
04_py_openai_compatible_callpython_automationsyntax:python, must_include(3), must_not_include(2)1-5
05_py_factorial_constraintsinstruction_followingsyntax:python, must_include(1), must_not_include(2)1-5
06_py_retry_backoff_decoratorpython_automationsyntax:python, must_include(2), must_include_any(3)1-5
07_py_concurrent_endpoint_pingpython_automationsyntax:python, must_include(2), must_include_any(3)1-5
08_py_parse_nvidia_smi_powerpython_automationsyntax:python, must_include(3)1-5
09_py_phone_regexpython_automationsyntax:python, must_include(2), must_include_any(2)1-5
10_systemd_timer_oncalendarconfig_editmust_include(3), must_not_include(1)1-5
11_yaml_add_fleet_tierconfig_editmust_include(5)1-5
12_htaccess_301_httpsconfig_editmust_include(4), must_include_any(3)1-5
13_php_canonical_tagpython_automationmust_include(2), must_include_any(2)1-5
14_sql_top_referring_domainssqlmust_include(5), must_include_any(4)1-5
15_debug_keyerrordebug_fixsyntax:python, must_include(1), must_include_any(4)1-5
16_debug_ollama_cpu_onlydebug_fixmust_include_any(6)1-5
17_refactor_bash_looprefactorsyntax:bash, must_include(3), must_not_include(1)1-5
18_refactor_preserve_signatureinstruction_followingsyntax:python, must_include(1)1-5
19_git_branch_commit_pushinstruction_followingsyntax:bash, must_include(3), must_include_any(2)1-5
20_py_argparse_clipython_automationsyntax:python, must_include(4)1-5
21_multifile_config_not_loadedmulti_file_reasoningsyntax:python, must_include(2), must_include_any(3)1-5
22_py_idempotent_insert_guardpython_automationsyntax:python, must_include(2), must_include_any(4)1-5

Coding v2 - hard set (20 cases, added 2026-07-05)

ModelScoreRuns (range)PassLatencyTempMeasured
Claude Opus-4.8 (cloud reference) BEST0.974n=120/207.7sdefault2026-07-06
Ornith-1.0-397B (FP8)0.971n=3 (0.966-0.977)20/2013.0sdefault2026-07-05
GLM-5.2 FP8 (Z.ai serving, reference)0.952n=3 (0.927-0.965)19/2078.5sdefault2026-07-06
GLM-5.2 (NVFP4)0.942n=3 (0.918-0.968)20/20147.6sdefault2026-07-06
Qwen3-Coder-Next-80B0.939n=3 (0.930-0.950)20/203.0sdefault2026-07-05
GPT-5.4 (cloud reference)0.935n=120/204.0sdefault2026-07-06
MiniMax-M3 (MXFP8)0.931n=3 (0.923-0.936)18/2030.0sdefault2026-07-06
GPT-5.2-Codex (cloud reference)0.928n=119/207.1sdefault2026-07-06
MiniMax-M2.7 (NVFP4)0.901n=119/2048.8sdefault2026-07-05
Qwen3.6-35B-A3B0.892n=3 (0.879-0.903)18/2027.4sdefault2026-07-06
The original 22-case suite saturated (the top ten sit within 0.02, below the noise floor), so v2 was built to discriminate at the top: multi-file bug tracing (import cycles, config precedence, systemd environment, cache-header rollouts), hard debugging (mutable defaults, DST arithmetic, subprocess deadlock, asyncio fan-out), strict spec-compliance tasks where any missed constraint costs points, gaps-and-islands SQL, and infrastructure tasks (Quadlet units, hotfix git surgery, parallel-safe bash). Same scoring blend and GPT-4.1 judge as v1. The hard set separates models the saturated suite could not: leaders that tie at 0.97-0.99 on v1 spread across 0.89-0.97 here. v1 remains the longitudinal reference; v2 decides picks.
Per-case breakdown: coding_tasks_v2.jsonl (20 cases; full prompt texts withheld to keep the benchmark uncontaminated)
Case idCategoryAutomatic checksJudge scale
v2_01_interval_merge_edgealgorithmssyntax:python, must_include(3), must_include_any(2)1-5
v2_02_toposort_cyclealgorithmssyntax:python, must_include(4), must_include_any(3)1-5
v2_03_lru_ttlalgorithmssyntax:python, must_include(6), must_not_include(2)1-5
v2_04_debug_mutable_defaulthard_debugsyntax:python, must_include(2), must_include_any(4), must_not_include(1)1-5
v2_05_debug_tz_dsthard_debugsyntax:python, must_include(2), must_include_any(4), must_not_include(1)1-5
v2_06_debug_subprocess_deadlockhard_debugsyntax:python, must_include(4), must_not_include(1)1-5
v2_07_debug_async_gatherhard_debugsyntax:python, must_include(2), must_include_any(2)1-5
v2_08_spec_versioned_configspec_compliancesyntax:python, must_include(4), must_not_include(3)1-5
v2_09_spec_redact_loggerspec_compliancesyntax:python, must_include(6), must_include_any(2)1-5
v2_10_spec_atomic_writespec_compliancesyntax:python, must_include(5), must_not_include(2)1-5
v2_11_spec_cli_exitcodesspec_compliancesyntax:python, must_include(4), must_include_any(2), must_not_include(2)1-5
v2_12_multifile_import_cyclemulti_file_reasoningsyntax:python, must_include(3), must_include_any(3)1-5
v2_13_multifile_env_precedencemulti_file_reasoningsyntax:python, must_include(2), must_include_any(4)1-5
v2_14_multifile_systemd_envmulti_file_reasoningmust_include(3), must_include_any(4)1-5
v2_15_multifile_nginx_cachemulti_file_reasoningmust_include(3), must_include_any(4)1-5
v2_16_sql_window_dedupsql_hardmust_include(3), must_include_any(3)1-5
v2_17_sql_upsert_countersql_hardmust_include(5), must_include_any(2)1-5
v2_18_bash_parallel_safeinfra_hardsyntax:bash, must_include(4), must_include_any(4)1-5
v2_19_podman_quadletinfra_hardmust_include(7), must_include_any(2)1-5
v2_20_gitops_hotfixinfra_hardsyntax:bash, must_include(6), must_include_any(3)1-5

Research & reasoning (27-case, none-context, task-appropriate temp=0.0-0.3)

ModelScoreRuns (range)PassLatencyTempMeasured
Qwen3.6-27B (dense) BEST0.963n=126/2737.6sdefault2026-06-30
Qwen3-Coder-Next-80B0.963n=126/272.0sdefault2026-07-05
Gemma-4-31B (BF16)0.963n=126/2711.7sdefault2026-06-09
Gemma-4-26B-A4B (BF16)0.963n=126/272.2sdefault2026-06-09
Nex-N2-Pro (NVFP4)0.963n=126/273.2sdefault2026-06-10
DeepSeek-V4-Flash-DSpark (2x GB10)0.963n=126/2725.9sdefault2026-06-30
Ornith-1.0-397B (FP8)0.963n=126/2722.9sdefault2026-07-05
MiniMax-M3 (MXFP8)0.963n=126/2715.6sdefault2026-07-05
gemma-4-12b-nvfp40.963n=126/275.0sdefault2026-07-15
gemma-4-26b-a4b-nvfp40.963n=126/273.5sdefault2026-07-15
DeepSeek-V4-Flash0.926n=125/2715.1sdefault2026-05-27
Claude Opus-4.70.926n=125/279.5sdefault2026-05-25
Qwen3.6-27B BF16 (dense)0.926n=125/27161.2sdefault2026-06-09
Gemma-4-12B (BF16)0.926n=125/275.6sdefault2026-06-09
Gemma-4-e4b (BF16)0.926n=125/272.7sdefault2026-06-09
GLM-5.2 (NVFP4)0.926n=125/2746.1sdefault2026-07-05
GLM-5.2 FP8 (Z.ai serving, reference)0.926n=125/2735.0sdefault2026-07-06
Qwen3.8-27B (FP8)0.926n=125/27719.9sdefault2026-08-14
Granite-4.1-8B0.889n=124/272.6sdefault2026-06-09
Qwen3.6-35B-A3B (BF16)0.889n=124/2724.2sdefault2026-06-08
devstral-small-2-24b0.889n=124/273.1sdefault2026-07-03
Gemma-4-e2b0.852n=123/2712.2sdefault2026-05-26
Gemma-4-31B0.852n=123/2743.2sdefault2026-05-27
Qwen3.5-Opus-distill (27B)0.852n=123/27270.5sdefault2026-05-27
Ministral-3-8B0.852n=123/2711.5sdefault2026-05-27
Ornith-1.0-35B (FP8)0.852n=123/278.8sdefault2026-07-02
Ornith-1.0-397B (W4A16 int4)0.852n=123/2726.5sdefault2026-07-05
MiniMax-M2.7 (NVFP4)0.852n=123/2712.2sdefault2026-07-05
Qwen3.6-35B-A3B0.815n=122/271.7sdefault2026-06-02
GLM-4.5-Air0.815n=122/27121.8sdefault2026-05-28
Qwen3-Coder-30B0.815n=122/271.7sdefault2026-06-09
Muse-Glimmer-30B (Meta)0.815n=122/2733.8sdefault2026-08-15
Gemma-4-e4b0.778n=121/2720.0sdefault2026-05-26
North-Mini-Code-1.0 (FP8, Cohere)0.778n=121/276.9sdefault2026-06-20
nemotron-3.5-lightning-30b-a3b-nvfp40.778n=121/274.8sdefault2026-08-11
Mistral-Small-4-119B0.741n=120/2711.0sdefault2026-05-26
Mistral-Small-24B0.741n=120/274.2sdefault2026-06-09
Gemma-4-26B-A4B0.556n=115/2746.4sdefault2026-05-26
Heretic-9B0.000n=10/27-default2026-05-27
Re-baselined on a single cloud judge (Haiku-4.5 or GPT-4.1) per Session-13 lock-in. Qwen3.6-35B-A3B, DeepSeek-V4-Flash and Claude Opus-4.7 tie at 0.926 - fleet reasoning is at Opus parity. The Claude-Opus reasoning-distill (Qwen3.5-Opus, 0.852) did NOT beat native dense.
Per-case breakdown: research_tasks.jsonl (27 cases; full prompt texts withheld to keep the benchmark uncontaminated)
Case idCategoryAutomatic checksJudge scale
summ_01long_doc_summarizationmust_include(4), must_not_include(3)1-5
summ_02long_doc_summarizationmust_include(3), must_not_include(2)1-5
summ_03long_doc_summarizationmust_include(4), must_not_include(2)1-5
summ_04long_doc_summarizationmust_include(3), must_not_include(2)1-5
summ_05long_doc_summarizationmust_include(4), must_not_include(2)1-5
summ_06long_doc_summarizationmust_include(3), must_not_include(2)1-5
summ_07long_doc_summarizationmust_include(7), must_not_include(2)1-5
mhqa_01multi_hop_qamust_include(3), must_include_any(3), must_not_include(2)1-5
mhqa_02multi_hop_qamust_include(2), must_not_include(4)1-5
mhqa_03multi_hop_qamust_include(2), must_include_any(5), must_not_include(2)1-5
mhqa_04multi_hop_qamust_include(3), must_include_any(6)1-5
mhqa_05multi_hop_qamust_include(2), must_include_any(3), must_not_include(1)1-5
mhqa_06multi_hop_qamust_include_any(7), must_not_include(2)1-5
mhqa_07multi_hop_qamust_include(2), must_include_any(4)1-5
cite_01citation_accuracymust_include(3), must_include_any(2), must_not_include(4)1-5
cite_02citation_accuracymust_include_any(6), must_not_include(4)1-5
cite_03citation_accuracymust_include(1), must_include_any(2), must_not_include(3)1-5
cite_04citation_accuracymust_include_any(7), must_not_include(4)1-5
cite_05citation_accuracymust_include(4), must_include_any(2), must_not_include(3)1-5
cite_06citation_accuracymust_include_any(9), must_not_include(4)1-5
code_01code_generationsyntax:bash, must_include(5), must_include_any(2), must_not_include(1)1-5
code_02code_generationsyntax:python, must_include(6), must_include_any(2), must_include_all(3)1-5
code_03code_generationsyntax:fish, must_include(5), must_include_any(2), must_not_include(2)1-5
code_04code_generationsyntax:python, must_include(10), must_include_any(3), must_not_include(4)1-5
code_05code_generationsyntax:bash, must_include(5), must_include_any(2), must_not_include(2)1-5
code_06code_generationsyntax:python, must_include(7), must_include_any(2), must_not_include(1)1-5
code_07code_generationsyntax:bash, must_include(10), must_include_any(2), must_not_include(1)1-5

Blog writing (10-criteria, combined = 0.5 structural + 0.5 judge, best-per-model temp)

ModelScoreRuns (range)WordsJudgeGen speedTempMeasured
DeepSeek-V4-Flash BEST1.000n=23774-64s 0.32026-05-28
Qwen3.6-27B BF16 (dense)1.000n=1382910/10477s 28 t/sdefault2026-06-08
Nex-N2-Pro (NVFP4)1.000n=1441110/1095s 91 t/sdefault2026-06-10
Qwen3.6-27B (dense)1.000n=1330510/1081s 98 t/sdefault2026-06-30
qwen3.6-27b-bf161.000n=3 (1.000-1.000)315710/10148s 60 t/sdefault2026-07-02
Ornith-1.0-397B (FP8)1.000n=3 (1.000-1.000)385710/1064s 120 t/sdefault2026-07-05
GLM-5.2 (NVFP4)1.000n=1343910/10207s 40 t/sdefault2026-07-06
MiniMax-M3 (MXFP8)1.000n=3 (1.000-1.000)331710/1046s 123 t/sdefault2026-07-05
Claude Opus-4.8 (cloud reference)1.000n=1329710/10125s 77 t/sdefault2026-07-06
Claude Fable 5 (cloud reference)1.000n=1363010/10135s 81 t/sdefault2026-07-06
Ornith-1.0-35B (FP8)0.963n=3 (0.945-1.000)221010/1066s 166 t/sdefault2026-07-02
gpt-oss-120b0.945n=21938-132s 0.02026-05-28
Gemma-4-31B (BF16)0.945n=1239010/10169s 24 t/sdefault2026-06-09
Gemma-4-12B (BF16)0.945n=1229110/1078s 53 t/sdefault2026-06-09
North-Mini-Code-1.0 (FP8, Cohere)0.945n=1226910/1043s 139 t/sdefault2026-06-20
DeepSeek-V4-Flash-DSpark (2x GB10)0.945n=1264610/10131s 46 t/sdefault2026-06-30
Laguna-S-2.1-NVFP40.945n=1325510/1055s 106 t/sdefault2026-07-26
Ornith-1.0-397B (W4A16 int4)0.945n=3 (0.944-0.945)32519/1070s 104 t/sdefault2026-07-03
GLM-4.7-Flash0.945n=22640-59s 0.02026-05-28
Qwen3.6-35B-A3B (BF16)0.944n=132939/1049s 171 t/sdefault2026-06-08
GLM-5.2 (IQ2_M 2-bit)0.944n=141279/10199s 39 t/sdefault2026-06-20
GPT-5.4 (cloud reference)0.944n=132469/1064s 102 t/sdefault2026-07-06
MiniMax-M2.7 (NVFP4)0.907n=3 (0.833-1.000)39669/1050s 129 t/sdefault2026-07-05
Qwen3.6-35B-A3B (FP8)0.889n=22597-38s 0.02026-05-28
Gemma-4-26B-A4B0.889n=122809/10128s 44 t/sdefault2026-05-26
Gemma-4-31B0.889n=120639/10553s 7 t/sdefault2026-05-27
Gemma-4-26B-A4B (BF16)0.889n=1212810/1026s 147 t/sdefault2026-06-09
Gemma-4-e4b (BF16)0.889n=1267610/1040s 122 t/sdefault2026-06-09
Laguna-XS-2.1-NVFP40.889n=117039/1030s 192 t/sdefault2026-07-26
Qwen3-Coder-Next-80B0.870n=3 (0.833-0.945)32109/1036s 180 t/sdefault2026-07-05
gemma-4-26b-a4b-nvfp40.870n=3 (0.833-0.945)18279/1034s 97 t/sdefault2026-07-15
Qwen3.6-27B-NVFP4 (unsloth, dense)0.855n=5 (0.333-1.000)32119/1089s 98 t/sdefault2026-07-12
Gemma-4-e2b0.833n=128649/1080s 78 t/sdefault2026-05-26
command-a-plus0.833n=119539/10103s 49 t/sdefault2026-05-26
Qwen3.5-Opus-distill (27B)0.833n=141979/10240s 37 t/sdefault2026-05-26
Laguna-M.1-NVFP40.833n=119269/1058s 79 t/sdefault2026-07-26
nemotron-3.5-lightning-30b-a3b-nvfp40.833n=135829/10104s 73 t/sdefault2026-08-11
Claude Opus-4.70.805n=22342-107s default (~0.0)2026-05-28
Qwen3.6-35B-A3B-NVFP4-Fast (unsloth)0.800n=5 (0.278-0.945)271910/1056s 138 t/sdefault2026-07-12
devstral-small-2-24b0.796n=3 (0.778-0.833)13079/1032s 124 t/sdefault2026-07-03
gpt-oss-20b0.778n=117939/1034s 243 t/sdefault2026-05-25
GLM-4-9B0.778n=18987/1012s 181 t/sdefault2026-05-25
Mixtral-8x22B0.778n=19317/1041s 60 t/sdefault2026-05-26
Granite-4.1-8B0.778n=18689/1013s 169 t/sdefault2026-06-09
Qwen3.8-27B (FP8)0.778n=125228/10700s 8 t/sdefault2026-08-14
Mistral-Small-24B0.722n=113427/1039s 74 t/sdefault2026-06-09
Gemma-4-e4b0.722n=124347/10105s 47 t/sdefault2026-05-26
Mistral-Small-4-119B0.722n=132628/10281s 28 t/sdefault2026-05-26
gemma-4-12b-nvfp40.704n=3 (0.333-0.889)61383/10202s 59 t/sdefault2026-07-15
Mistral-Medium-3.5-128B0.611n=114697/10450s 18 t/sdefault2026-05-26
GLM-4.5-Air0.500n=145413/101360s 6 t/sdefault2026-05-25
nemotron-3-nano:30b0.445n=156182/10156s 51 t/sdefault2026-05-25
Qwen3.6-35B-A3B0.278n=11953/10212s 38 t/sdefault2026-05-25
Mixtral-8x7B0.222n=1792/101s 137 t/sdefault2026-05-26
Qwen3-Coder-30B0.111n=101/1011s 0 t/sdefault2026-06-09
3,000+ word SEO CMMC blog from a fixed prompt. Each row reports the model's best-scoring temperature when the Session-14 Bench-2 sweep covered it (the 2026-05-28 model set only); models benched after that date show their single measured temperature. Scores within 0.03 are statistical ties. Autoblog runs overnight in batch, so quality decides.

Blog temperature sweep (Session-14 Bench-2, judge=GPT-4.1, N=2 reruns per cell)

ModelTempMean scoreRange (min-max)Mean wordsMean gen
Claude Opus-4.7default (~0.0)0.8050.778-0.8332342107.2s
Claude Opus-4.7default (~0.3)0.8050.778-0.8332342107.2s
Claude Opus-4.7default (~0.7)0.8050.778-0.8332342107.2s
Claude Opus-4.7default (~1.0)0.8050.778-0.8332342107.2s
DeepSeek-V4-Flash0.00.9720.944-1.000419478.4s
DeepSeek-V4-Flash0.3 BEST1.0001.000-1.000377464.2s
DeepSeek-V4-Flash0.71.0001.000-1.000373265.0s
DeepSeek-V4-Flash1.00.9720.944-1.000402572.4s
GLM-4.7-Flash0.0 BEST0.9450.944-0.945264058.8s
GLM-4.7-Flash0.30.8050.722-0.889221053.1s
GLM-4.7-Flash0.70.8610.833-0.889274857.9s
GLM-4.7-Flash1.00.8890.889-0.889207249.5s
Qwen3.6-35B-A3B (FP8)0.0 BEST0.8890.889-0.889259737.5s
Qwen3.6-35B-A3B (FP8)0.30.8060.667-0.945320439.8s
Qwen3.6-35B-A3B (FP8)0.70.5840.445-0.722251744.2s
Qwen3.6-35B-A3B (FP8)1.00.6110.611-0.611254941.5s
gpt-oss-120b0.0 BEST0.9450.945-0.9451938131.8s
gpt-oss-120b0.30.8890.889-0.8892145137.4s
gpt-oss-120b0.70.9450.945-0.9452144127.8s
gpt-oss-120b1.00.9170.889-0.9451710124.2s
Session-14 Bench-2 temperature sweep measured 2026-05-28 13:48 EDT. Reasoning models (DeepSeek-V4-Flash, Qwen3.6-35B-A3B) may treat temperature as a hint during their reasoning phase; a flat row across temps is itself a finding. Claude Opus-4.7 rejects the temperature parameter; its rows are copied from a single default-temperature run.

How the scores are produced

Coding - 22 real operations tasks, not textbook puzzles. Drawn from Petronella Technology Group's day-to-day infrastructure and SEO work: 8 Python automation tasks (OpenAI-compatible API clients, retry/backoff decorators, concurrent endpoint health checks, log parsing, argparse CLIs), 3 bash/fish shell-ops tasks (timestamped backups, safe rsync deploys, fleet uptime sweeps), 3 production config edits (systemd OnCalendar timers, YAML inventories, .htaccess 301 rules), 3 instruction-following traps (tasks that fail if a stated constraint is ignored, such as preserving a function signature), 2 debug-and-fix cases, 1 SQL analytics query, 1 shell refactor, and 1 multi-file reasoning case. Each case is scored 0.6 x automatic checks (required content, regex gates, and real syntax validation: bash -n, python compile, fish -n) + 0.4 x a blind LM-judge correctness score (1-5; the judge sees only task and solution, never the model's identity). Deterministic decoding (temp 0.0). The leaderboard number is the mean across all 22 cases.
Research & reasoning - 27 cases in four categories. 7 long-document summarization cases (word bounds, bullet counts, must-include / must-not-include gates against supplied source documents), 7 multi-hop QA cases (the answer requires chaining facts across documents), 6 citation-grounding cases (presence and format of citations regex-verified against the supplied sources; semantic citation correctness is measured separately in PTG's cited-RAG evaluations, which score a much harsher 0.65-0.80 F1), and 7 grounded code-generation cases with syntax gates. Same 0.6 auto + 0.4 blind-judge blend as coding.
Blog writing - one fixed brief, structural gates plus publish-readiness. Every model gets the identical brief: a 3,000+ word CMMC compliance post in clean HTML. Nine automatic structural gates (word count, H1/H2/H3 hierarchy with 8+ sections, no leftover markdown artifacts, FAQ with 5+ substantive Q&As, a well-formed comparison table, internal links, correct call-to-action, and brand rules) plus a 10-criteria publish-readiness judge score (1-10) covering factual accuracy (CMMC 2.0 levels, NIST 800-171), tone, and SEO structure. Combined = 0.5 x structural pass rate + 0.5 x judge. Blog rows report each model's best temperature from the Session-14 sweep where covered.
Voice tool-calling - 88-case corpus with decline-safety distractors. Real assistant tool schemas; the corpus mixes valid calls with hard distractors that look like tool calls but must be refused. Scored on correct calls, correct declines, and hallucinated-call rate. This suite decides the voice tool-caller pick; decline-safety on distractors is the deciding metric.
Agentic coding - SWE-bench Verified, no LM judge. Stratified 150-instance subset of SWE-bench Verified run through the official harness with mini-swe-agent; a task counts only if the generated patch resolves the issue's test suite. Referenced in picks and findings (for example DeepSeek-V4-Flash 67.3%, measured 2026-06-28); not a leaderboard column.
Throughput - measured on the serving hardware. Single-stream (c=1) plus concurrent c=8 and c=32 sweeps of 800-token generations against the exact serve config named in each finding. Speeds quoted in findings are decode tokens per second on the stated GPUs and quant.
Judging integrity - the cross-judge gate. A generator never grades itself: self-judging inflated scores 20-37% in our measurements. The campaign judge is locked to GPT-4.1 (code-enforced since 2026-05-30, with a guard that flags any same-family judge pairing). Legacy rows graded by Claude Haiku-4.5 are labeled; DeepSeek-V4-Flash is accepted for bulk grading after validating equivalence with GPT-4.1 (r=0.962, mean absolute delta 0.18 on the 1-5 scale; re-audited 2026-07-06 on 40 blog outputs: mean abs delta 0.45 on the 1-10 scale, 87.5% within 1 point, second judge slightly stricter). Excluded from all tables: misconfigured or errored runs, runs without a cross-family judge, and hardware-specific Mac coding runs (see the Apple Silicon matrix).
Reproducibility. Every row carries an absolute measurement date; the exact quantization is recorded in the run label (quant changes the result: the same model at FP8, NVFP4, and int4 scores differently and is listed separately). Headline claims run N=3 to N=5 repetitions. As of 2026-07-05 the board shows the MEAN of each model's largest rep family (newest on ties) with n and min-max range in the Runs column (previously best-run, which rewarded models benched more often); scores within 0.03 are statistical ties.

Findings

Measured 2026-07-15
Is Gemma-4-26B-A4B better than Qwen3.6, and can it run in the agent gateway? It depends on task difficulty, judged by a three-model frontier panel (GPT-5.4 + Claude Opus-4.8 + GLM-5.2, median score on identical generations, N=5). On easy/routine coding Qwen wins (dense 27B 0.983, 35B-FP8 0.979 vs Gemma 0.962); on the hard coding set Gemma wins decisively (0.915 vs every Qwen variant 0.78-0.81), corroborated by the deterministic auto-checks. Research and blog are ties (about 0.86 and 0.79). Agentic: served with the correct vLLM tool parser (gemma4, not pythonic - the wrong one produces zero tool calls), Gemma is the strongest tool-caller (55 correct calls vs Qwen's 51) and solves every multi-turn task (pass@max 1.00 vs Qwen-FP8 0.778). Both models were verified end-to-end through the real agent gateway - writing and executing code and returning correct results - which requires serving the model at 131k context (the agent system prompt is about 58k tokens). NVFP4 is near-lossless for Gemma-4 but not for Qwen3.6-35B (a long-form runaway appears at FP4; use FP8).
Measured 2026-07-06
Does GLM-5.2 lose anything at FP4, and is the six-GPU FP8 layout justified? No measurable loss. GLM-5.2 at FULL FP8 fidelity (Z.ai's own serving, via OpenRouter) scored 0.952 on the v2 hard set (n=3: 0.927/0.964/0.965) and 0.92 on research - versus the FP4 quant on the fleet's 4-way bridge at 0.942 and 0.94. The FP8 delta (+0.010 coding, -0.02 research) is inside the 0.03 tie threshold: FP4 is effectively lossless for GLM-5.2 on these tests, and both configurations score below Ornith-1.0-397B-FP8 (0.971). Implication for fleet layout: an FP8 deployment spanning six GPUs and both NVLink bridges buys no measurable quality over FP4 on four GPUs.
Measured 2026-07-06
Do two independent judges agree on these scores? Yes. A dual-judge audit re-graded 40 stored blog outputs with DeepSeek-V4-Flash (served on the DGX Spark pair) against their original GPT-4.1 scores using the identical rubric: mean absolute difference 0.45 points on the 1-10 scale, 87.5% of pairs within one point, and the second judge ran slightly STRICTER on average (signed mean -0.3) with its five largest disagreements spread across Anthropic, MiniMax, and Qwen outputs - no family favoritism pattern. This is the periodic cross-judge integrity check the methodology commits to.
Measured 2026-07-06
How does the fleet compare to the July 2026 cloud frontier on the same tests? Cloud reference rows added: Claude Opus-4.8 v1 0.995 / v2 0.974 / blog 1.000 (GPT-4.1 judge); GPT-5.4 v1 0.970 / v2 0.935 / blog 0.944 and GPT-5.2-Codex v1 0.982 / v2 0.928 (both Haiku-judged per the cross-family gate; OpenAI rows and Anthropic/fleet rows use different judges, so treat cross-vendor deltas under 0.05 with care). Read: fleet-owned Ornith-1.0-397B-FP8 (v2 0.971, zero marginal cost, CUI-capable) sits between GPT-5.4 (0.935) and Opus-4.8 (0.974) on the hard set. Also measured: Claude Fable 5's API safety layer CONTENT-FILTERED 9 of 22 routine infrastructure tasks (finish_reason=content_filter after ~3 tokens: rsync deploy scripts, a phone-number regex, an argparse CLI), making it unusable as a coding baseline here; its blog run scored 1.000. GPT-5.2-Codex is Responses-API-only (the harness gained --use-responses-api).
Measured 2026-07-06
Do the bridge picks survive a harder test set and multi-topic writing? Yes, and the ordering sharpens. The 20-case v2 coding set (built because the original suite saturated) spreads the leaders decisively: Ornith-1.0-397B-FP8 0.971 (n=3, 20/20 every rep, 13 s per case), GLM-5.2-NVFP4 0.942 (n=3, high variance 0.918-0.968 at 150 s per case), Qwen3-Coder-Next 0.939 (n=3, 2.6 s per case), MiniMax-M3 0.931 (n=3, drops 1-2 cases per rep), incumbent Qwen3.6-35B-A3B 0.892 (n=3). MiniMax-M3's v1 crown (0.984) inverts on v2: the saturated suite was measuring judge quibbles, not capability. Blog re-tested at 3 topics x 3 reps: Ornith-397B a perfect 1.000 on all nine runs; MiniMax-M3 0.975 mean (short only on the CMMC brief); MiniMax-M2.7 0.94 mean (its earlier single-run 1.000 was variance); Qwen3-Coder-Next 0.887 mean. Qwen3-Coder-Next also posted 85.2% on the 88-case tool corpus (53 correct calls, 24 correct declines, 1 hallucinated call): strong for agentic coding, not voice-grade (gpt-oss-20b remains 0% hallucination). Remaining before any production routing change: a local SWE-bench Verified run for Qwen3-Coder-Next.
Measured 2026-07-05
Which open model is best on each NVLink island of the 6x H200 fleet? First head-to-head since the 4-way bridge install. On the 4-way bridge (575 GB), Ornith-1.0-397B-FP8 (MIT) is the balanced winner: coding 0.980 (N=3), blog 1.000, research 0.91 at 119.7 t/s single-stream and 639 t/s at c=8, roughly 2x the speed of speed-tuned GLM-5.2-NVFP4 at equal or better quality on two of three axes. MiniMax-M3-MXFP8 posted the campaign's best coding rep (0.991; N=3 mean 0.984) with 1M context and multimodal input, held back only by its non-MIT community license. GLM-5.2-NVFP4 keeps the research crown (0.94, citation accuracy 0.992). On the 2-way bridge (287 GB), Qwen3-Coder-Next-FP8 (80B-A3B) is the coder/agentic pick: 0.977 coding with zero variance across three full reps, clean native tool calls, 178 t/s single-stream and 2,789 t/s aggregate at c=32. Ruled out: Kimi-K2.6 (594 GB INT4 exceeds any island), GLM-5.x-FP8 (754 GB needs all six cards), Nemotron-3-Ultra-NVFP4 (numerically broken on Hopper: loads but emits gibberish). Serving notes: Ornith-397B-FP8 requires VLLM_TEST_FORCE_FP8_MARLIN=1 plus an explicit chat template; MiniMax-M3 requires --block-size 128; the 2026-07-03 vLLM nightly has broken block-scaled FP8 kernels on Hopper (pin v0.24.0).
Which models do we keep resident and route to? Coding → Qwen3.6-35B-A3B (0.975, GPT-4.1). Research → Qwen3.6-35B (speed) / DeepSeek-V4-Flash (long-context), both Opus-parity. Blog → DeepSeek-V4-Flash / GLM-4.7-Flash. Voice tool-call → gpt-oss-20b. Edge appliance → Granite-4.1-8B (+ Gemma-4-e4b for tool-driving). Reserve Claude Opus for the top few percent.
Do self-improving loops help small models? A loop is a capability amplifier, not an equalizer: Qwen3.6-35B goes 0.917→1.0 with iterations; small models (Gemma-4-e4b, GLM-4.7-Flash) stay flat at 0.917. The done-gate makes a small model honest (no silent fabrication), not capable.
Self-improving loop vs a general agent (pi.dev) on a 4B? Gemma-4-e4b scored 0.917 in a done-gated loop vs 0.0 in pi.dev (it fabricated all 12 tasks). Air-gap appliances should pair a small model with a programmatic verifier loop, never a general agent framework.
Are there other open models worth adding? A live Feb–May 2026 scan found none that beat the incumbents; Qwen3.7/Qwen4, DeepSeek-R2, Phi-5 and Grok-3 weights are unreleased or hosted-only.
What is the best model for a single RTX PRO 6000 96GB (Blackwell) card, and is a 35B a waste of it? No displacer. Qwen3.6-35B-A3B (3B active) is the best all-rounder that fits one card: it wins research and cited-RAG outright and leads coding (0.975, GPT-4.1). The models large enough to “fill” the card (gpt-oss-120b 0.964, Mistral-Small-4-119B 0.957) are slower and weaker on the role’s core axes. A low-active MoE is the correct shape for a 96GB concurrency server: comparable NVFP4 models scale to ~2,000 t/s aggregate at c=32 on this card. Spare VRAM is best spent on KV/concurrency, or on NVFP4 (same quality at half the VRAM, freeing room to co-locate a second model), not on a bigger-but-worse model. Mistral-Small-4-119B is the lone alternative, and only if the card is redefined as a cited-RAG / vision / compliance resident.
Is a purpose-built Rust inference engine (Atlas) faster than our tuned vLLM on Blackwell? (measured 2026-06-07) No. On identical GB10 (DGX-Spark-class) hardware and the same Qwen3.6-35B-A3B-NVFP4 model, our tuned vLLM (NVIDIA MTP recipe) ran 116–119 tok/s steady-state vs Atlas’s 88.9; Atlas’s advertised “130–133 tok/s” and “3.1× faster than vLLM” did not reproduce (the 3.1× is vs an untuned vLLM). Atlas serving is quality-preserving - blog 0.944 (ties our blog leader) and 6/6 on a coding spot-check - and ships an ~8×-smaller (2.98 GB) no-Python single binary. That makes it a candidate packaging vehicle for an air-gapped compliance appliance, not a throughput upgrade. Its multi-node expert-parallel mode is not yet shipping (runtime is single-node only).
Does an agentic multi-hop retriever beat single-shot RAG for compliance Q&A? (measured 2026-06-07) On hard multi-hop CMMC / NIST 800-171 questions, an RL-trained search agent (Harness-1, 21B, gpt-oss-20b base) found every gold control (retrieval recall 1.000) where single-shot dense top-8 reached only 0.881 - it recovers the deep 2nd/3rd-hop controls single-shot drops at production cutoffs. But its curated answer (0.929) only matched single-shot top-15 (the curation step, not the search, is the bottleneck) and cost ~1,000× the latency - so the value is exhaustive batch retrieval (audit / SSP gap analysis), not interactive RAG. Control-id deduplication remains the cheap universal lever: it lifts both single-shot (0.786→0.881) and the agent (0.905→1.000).
Benchmarked and published by Petronella Technology Group, Inc. All scores are produced by the PTG llm-benchmark harness described in "How the scores are produced" above.