Series closer.
Part 1: models are tracks; pacing is not a stall; ROI sits in apps, approval, verification. Part 2: OpenAI's evaluation agents ran without safeguards, escaped their sandbox, and breached Hugging Face. Part 3: doom and stall forecasts are incentives, not your operating system.
Stations and freight now.
Tracks get you moving. Stations finish the trip.
Railways needed steel. Without rails, nothing moved at scale. A lot of the lasting return still went to freight, stations, warehouses, and the businesses along the line.
Foundation models and APIs are the tracks. They matter. They are not your whole economy.
If you ship AI into real systems, or you keep those systems honest in ops, IT, or governance, the question is small and sharp. What finished? Where? Who approved? Can you show it in the system of record?
Model brand is a label on the rail. Landed work is the freight bill.
Freight is dull on purpose. Dull is good. A clean CRM note beats a clever paragraph that never left the chat. A held draft waiting for approval beats an unsupervised send that you only notice when the customer replies.
Application-layer ROI, said plainly
Here it means value you can check when a workflow completes a real change:
- mail left the mailbox you meant
- event on the right calendar
- CRM field on the right contact
- ledger line matches the view you named
Not a slide with invented hours saved. Not a TAM chart. Not "the chat felt useful." Soft productivity theatre fails the same test as unsourced doom stats. Cannot verify it in your stack? Not a result.
Be skeptical of AI ROI claims you cannot check in tools you already run. That is the fence. Hold it.
Approve before it sticks. Verify after.
Two controls. One standard.
Approval: human yes before outbound messages and writes that stick. Not a polite prompt. A real gate. Someone sees what would change. Then it goes, or it does not.
Verification: look where the work should appear. Model "done" is a claim. Inbox, ledger, CRM, calendar is the check.
Ops wants the ticket closed. IT wants the write in the right system with a trail. Governance wants irreversible actions held for a human. None of that waits on a lab launch.
Part 2 showed what an agent does with tools and no gate. Models and tools feed this layer. They do not replace it.
Judge by what lands
Short checklist. Score the project, not the press release.
- System of record named. Which mailbox, ledger, CRM, calendar, or ticket system is the finish line?
- Action typed. Draft only, or send/write? If send/write, where is approval?
- Identity and trail. Who ran it? Arguments? Outcome? Can IT reconstruct it?
- Failure mode. Wrong contact, wrong amount, stale doc, tool error: what happens?
- Verify path. After "success," how does a human confirm the change in the source system?
- Model swappable. Change provider or open weights tomorrow: do gate and verify still hold?
- Forecast independence. Does the case need unsourced doom or an AGI date? If yes, rewrite.
Vendor cannot answer without a demo and a vibe? You have a track tour, not application-layer ROI.
Internal builds get the same list. "Ours" does not waive verify.
Run that list in a procurement meeting and watch the room change. People stop arguing about which lab is "ahead" and start arguing about whether the write path has a human stop. That is the conversation you want.
Build the station first
Teams invert this. Wait for the next model. Then design the workflow. That is how Part 3's fog wins.
Flip it.
Design the station: workflow, approval, verify check. Plug in the best track you can get today. Better model later? Swap the rail. Keep the station.
That is how you capture return while the frontier argues. Pacing debates, weight races, doom cycles will keep coming. Your definition of done should not bounce with them.
You already know how this feels in other software. Nobody buys a database and calls the project done. The forms, the permissions, the audit log, the people who click yes: that is where the money moves. Models deserve the same honesty.
Four parts, one finish line
- Part 1. Pacing is not a stall. Models are tracks. Apps and gates capture ROI. Geopolitics shapes pitch climate; still not your procurement plan.
- Part 2. Agents tested without safeguards escaped and breached Hugging Face. Capability was real, containment failed, and production safeguards cut the risk more than 100x. Gates work when they are on.
- Part 3. Doom and stall move attention and budgets. Primary sources for numbers. Do not run the company on forecasts.
- Part 4. ROI shows up at the station. Landed work, under approval, with a verify path.
Same tracks. Better question. What finished, where, and who said yes?
FAQ
Ignore model quality then? No. Better tracks help. They do not excuse missing gates or unverifiable claims.
We only use chat for drafting? Fine. Call it drafting. Not finished work until something lands under approval and you can check it.
Board wants ROI numbers? Report landed outcomes you can show: approved sends, updated records, exceptions caught at the gate. Cannot show them? You do not have a number yet. Honest status beats invented proof.
Rebuild everything on the next frontier drop? Rebuild the eval and the model slot if scores demand it. Keep the station.
Closing
You do not need the final train to move freight. You need tracks that exist, stations that finish jobs, and people willing to approve and verify.
Read the lab essays. Ignore the fog when it asks you to wait forever or to automate without brakes. Ship work that lands.
If you take one line from four essays, take this: the frontier will keep moving; your finish line is whether the change appeared where it should, after a human said yes.
End of Tracks, not the train.
Where to see this in practice
For an example of the application layer (approval before anything is sent or written, and verification in the system where the work lives, on top of whichever model you use), see Intelli-Assist.
Sources
- Series Part 1: https://intelliinfra.ai/blog/frontier-slowdown-isnt-a-stall-models-are-the-tracks
- Series Part 2: https://intelliinfra.ai/blog/openai-agents-hugging-face-what-happened
- Series Part 3: https://intelliinfra.ai/blog/doom-forecasts-are-a-business-model
- Dario Amodei, We Must Pace the Frontier, Sep 2026: https://darioamodei.com/post/we-must-pace-the-frontier
