Frontier Inference Margins
An honesty-first interactive cost model of frontier-LLM inference serving margins — mechanism-first calculator, typed public-claims registry, and a long-form research report. Every number carries its epistemic status.
- Mechanism-first calculator — every assumption adjustable, each with its epistemic status attached
- 34-claim typed evidence registry with provenance tiers
- Read-only MCP connector at margins-mcp.ashitaorbis.com
- Three test suites and published release gates behind every deploy
Activity Timeline
-
Access fence protects superseded deployment hashes; 16 commits added ledger updates and a GPT-6 Astra estimate.
A Cloudflare Access fence now covers the project's Pages subdomains, protecting superseded deployment hashes. A publication retirement workflow for old release URLs was staged, and hourly dive-engine ticks produced 10 ledger updates.
-
Deployment correction shipped, but 78 of 81 retired deployment URLs still served stale releases 26+ minutes after deletion.
The scenario guard scope fix (commit 287b4307) was published after two review rounds on guarantee wording and path scope. Daily dc-map ledger and registry sweeps produced 26 automated commits tracking dive-engine state.
-
36 commits, 20 of them dive-engine ticks on the dc-map ledger and registry.
The sweep cursor advances only after a successful Grok collection. The Grok timeout rose to 1200 s, and a 7-day window cap was applied.
-
38 commits: share-card builder and template shipped to the public repo; release v3.0.0-2026-08-13 synced and published.
The agent-trajectory draft fold was integrated with the gate verdict and the sources table resolved with prior captures. Visual fixes keyed the rent bar, drew it to percent, and reordered the scenario window into reading order.
-
30 dive-engine heartbeat commits landed, and the 2026-09-26 daily sweep completed.
Continuous dc-map ledger and registry tracking ran throughout the day. The sweep finished with queue updates and GPT Pro material.
-
Hourly ledger and registry mapping ticks; Worker deployment updated.
Checkpoint fold moves and GPT Pro findings folded into README. Worker release 8dae23c0 re-minted.
-
41 commits from dive-engine automation; release candidate stamped and Worker artifact pinned.
Dive-engine ticked 20 ledger/registry update cycles across the day. GPT Pro checkpoint fold entered 6 prose findings into review with README phrases guarded.
-
Hourly dc-map ledger snapshots self-committing; review routing changed from Astra councils to single-pass.
Seven hourly snapshots committed between 09:30Z and 15:20Z. http_429 fix applied; late 2026-09-18 daily-material Pro report filed.
-
19-commit release live: same-assumption table bug fixed, cross-provider labels corrected, contract stamp conflict escalated.
Release addressed same-assumption table rendering and cross-provider relabeling across 19 commits. Contract registration conflict found during deployment: release stamp retargeted from 1092e5b to 2b480da. Conflict escalated for resolution.
-
Vetting mission completed; six corrections to equipment estimates, terminology, and cost reconstruction tables published across private and public mirror.
Corrections landed on master after multiple gate-verification rounds. Publisher credential hardening applied to Worker artifacts in the same release cycle.
-
41 commits folded E1–E5 and N1/P1 findings; release frozen at ccdca6c.
Core corrections retired an intermediate coefficient, fixed false attributions, resolved calibration record contradictions, and nailed down the E2 timing basis. Public mirror kept in sync across five mirroring commits.
-
Six prose repairs merged and synced: Trainium withdrawn, TPU/GB200 fixed, naming duplicates 35→0.
Outside-reading audit surfaced six issues; repairs covered precision contradictions, cost reconstruction, estimate vectors, and evidence label standardization. Five republish cycles propagated v3.0.0-2026-08-13 across inference-margins-public.
-
40-commit release: prose corrections, tariff updates, Zhipu fix, public mirror updated.
Six prose corrections landed alongside tariff updates and changelog revisions. Zhipu card fixed; intermediate coefficient retired from explanations tree-wide. UTC guard corrected and false attribution retracted from published materials.
-
17 commits: privacy gate closed across 5 surfaces, DC-map ledger index and sweep queue complete.
Privacy hardening addressed 9 code review findings and closed leakage paths in the ledger index, cache, render, rent-adjustment, and publish gate. DC-map pipeline received ledger index and sweep queue closure. Last deploy 07:44:56Z today.
-
8 commits fixing number discrepancies; 4 verifier logic bugs patched.
Corrected 63-vs-93 and 57-vs-93 reading discrepancies, glossary entries, attribution details, and re-mint notes. Four verifier defects fixed: guards, binding validation, adjudication logic, and false-comment bypass.
-
MCP HTTP server hardened against streaming edge cases; 3 new test suites passing.
Body-cap logic rewritten to pause/respond/destroy pattern; try/catch added to request handler; Content-Length precheck implemented. GPT Pro daily sweep completed, 2 commits applied findings.
-
Two minor corrections committed.
Unpushed commit count corrected from 31 to 35. China dossier western hub electricity routing claim corrected.
-
H800/H100 fit-transfer differential modeled as named adjustable assumption.
Daily sweep updated with Pro research outputs from 2026-08-18. Release stamp faf6bec deployed with served-asset manifest.
-
Desktop layout overhaul: page height 14,079 → 9,974 px, 7,933 assertions passed.
Result tile moved into projections band; estimates repositioned to desktop top; dossier rendering deduplicated to a single route and justification stack collapsed. Dual-viewport empirical verification: 7,933 PASS / 0 FAIL. 12 commits across 3 sessions.
-
Estimate-card faces moved to round-3 readings; analyst attribution and assumption disclosures restored.
Cards-vintage ruling enforced: estimate-card faces migrated from default to round-3. Two authorized MCP transport deltas declared for hoisted estimates block. Parity test re-minted for consistency verification.
-
Three stranded weekly Pro reports recovered; DeepSeek V4 tariff sweep queued.
Reports restored from 449-byte stubs via fixture recovery. Daily sweep tracking DeepSeek V4 peak/off-peak pricing (57–1,100% lifts) and X.ai/Z.ai positioning.
-
DeepSeek V4 tariff rewrite and Gemini 3.7 Flash launch documented.
DeepSeek V4 effective 2026-08-16: 2× peak/off-peak split, cache-hit markup increases up to +1,114%. Gemini 3.7 Flash launched 2026-08-13 at 50% introductory price cut. GPT Pro daily sweep dispatched; burn-guard reconciled harness packages and established 08-12 as canonical after sha256 check.
-
v3.0.0 released: 33 commits, GPT Pro review cleared, public mirror live.
Headline-invariance tests 255/255 clean, reverse-compatibility validated, migration-differential CI honesty and metrics discontinuity suppression fixed. GPT Pro returned NOT-READY (16 findings); all folded into bq-290–294 and verified clean on second pass.
-
Mean-based centroid for three-point sliders implemented and committed.
Slider semantics clarified: central value is MEAN not median for three-point mode. Centroid commit gates M8 chain. M4 restart sequenced after status report delivery.
-
Dual-consult enactment completed; NVIDIA-additive correction recorded.
Survivor-set exactness verified across both arms with declared ranges and reference vocabulary checked. GPT Pro arm landed with NVIDIA-additive correction; owner ruling applied. Max/min/median derived over provider mixes summing to 100.
-
Stack multiplier corrected to +2.4 months; v2.2.0 public sync.
1.25× frontier-lab assumption yields +2.4 months (was +2.9), post-Polaris adjudication. Public mirror synced to v2.2.0-2026-08-06. Contribution margin sweep: 77.31% engine median, 65–82% endpoint range.
-
v2.2.0 published to public mirror.
Public mirror updated to track private source hash 5a07e2f (2026-08-06 release).
-
Engine v2.1.12 deployed live; metric moved 77%→69% with documented decomposition.
13 commits shipped, public mirror synced. RAISE Summit podcast claims drove metric decomposition across size revision, evidence adjudication, and serving-model re-engineering. Post-deploy audit found 4 critical gaps: stale 200 OK responses, unverified hardening commits, missing post-deploy probes, and pending workers.dev disable.
-
DeepSeek V4 Flash: 35× cheaper than Kimi K3 on Vibe Code Bench.
Vals AI DeepSeek V4 Flash at $0.20/task vs Kimi K3's $17.56/task — material cost differential captured in daily sweep. workers.dev endpoints closed.
-
GPT-Pro fetcher patched; 07-23 through 07-27 daily material recovered.
Fetcher patched to recover the 07-27 collection gap. Previously missing data from 07-23/24/26 backfilled into the dataset. Automated daily sweep pipeline continues via async ChatGPT Pro dispatch.
-
Recovered truncated GPT Pro answers past a poisoned request_id cache.
Built recovery harness for three answers (28.8k–23.9k chars each). Poll loop with escalation logic now handles the timeout pattern. Cleared a sharp HIGH advisory in the worker dev tree.
-
CI hardened; GPT-Pro daily backlog through 2026-07-25 committed.
Read-only token, checksum-verified shfmt, and .pyc cleanup applied to CI. Daily material backlog for prior week persisted.
-
TPU7 +495.4% miss traced to operating-point-closure error; research pipeline automated.
Root cause: design inverted b*=154/chip vs actual ~16 concurrent-sequence ceiling. Standing recommendation: anchor future designs to 64-sequence batch ceiling. FlexNPU CM384 extracted for cycle-2. ChatGPT Pro polling now fully automated — dispatch, poll, persist.
-
Cycle 2 root cause found: 16 seq/chip actual vs ~154/chip assumed, explains +495.4% miss.
Operating-point-closure error traced and verified (TPU7 Ironwood: 64 concurrent sequences, 16/chip, 518.86 tok/s/chip). CM384 FlexNPU extracted via parallel research agents. Epoch AI and wafer_ai added to hardware sweep; GPT Pro replica-width consult refuted both prior candidate bounds.
-
IM3 exit verification in final rounds; IM4 diagnostic harness spec locked.
P0/P1 defect classes fixed and re-verified through Rounds 2–3. Script defect recovery advanced from r7 to r12 (r11 released defect-free). IM4 entry harness specified with survivorship regression fixtures.
-
ChatGPT Pro MCP disk_recovery premature-completion bug fixed; daily sweep on schedule.
Root cause of premature completion in disk_recovery path identified, documented, and patched in one commit. Daily automation pipeline persisting 2026-07-19 research findings (Z.ai, DeepSeek/Kimi, Anthropic credit line) to gptpro-reports/.
-
v2.2 verification in progress; deploy awaiting owner approval.
Ground-truth reconciliation pass underway against effective-MFU scalar model. ChatGPT Pro MCP completion-detection bug filed (commit 53b3dac). Deployment staged pending sign-off.
-
Pipeline hardened (4 P1 + 1 P2 findings applied); v2.2 on HOLD pending owner sign-off.
D2 receipt pack completed with full model/platform/operating-point verification. Cycle-2 re-engineering underway but threshold finalization blocked on owner gate.
-
v2.1.6→v2.1.11 (6 releases, 44 commits): billing bug fixed, GPT-Pro wired, cold review hygiene pass.
Chart billing canonicalized to shared computeMix engine with regression tests. Four dive reports scrubbed of provenance headers, analysis preserved. GPT-Pro daily/weekly reports live via bash-owned fetcher with in-turn polling and 70-min codex fallback.
-
v2.1.4: evidence-scent label pass, outside-review remediation, feedback intake live.
Page-set values relabeled so they no longer carry a false empirical scent; GPT-Pro outside-review findings remediated and the engine bumped to v2.1.4; quarantined feedback intake (Cloudflare Worker + D1 + Turnstile) embedded on the main page.
-
Margin-range evidence board redesign shipped; MCP connector deployed (v2.1.3).
Preset dropdowns replaced by a mechanism-first margin-range evidence board with de-named routes and v4 permalinks. Read-only MCP server deployed as a Cloudflare Worker at margins-mcp.ashitaorbis.com behind a fixture-gated honesty envelope.
-
Traffic-mix axis added to the engine (v2.1.2); reception audit round remediated.
New traffic-mix axis with a pure codec and atomic permalink replays, specified first by failing-by-design contract fixtures. Nine reception audits run against the live site; all 16 P0 findings remediated and the resolution ledger published.
-
Initial build and first production deploy.
Honesty-first interactive cost model of frontier-LLM inference serving margins: mechanism-first calculator, typed public-claims registry, and a long-form research report, deployed to Cloudflare Pages.