Almost 25 years ago I read an article on immobots in MIT Technology Review. It stuck: Immobots Take Control.
I was a fresh grad at Husky Injection Molding Systems, assigned to improve self-diagnostics on injection moulding machines and their robots. It was a lovely problem. It got me thinking about monitoring systems for errors — and, more importantly, about what an error is. An immobot doesn’t watch a machine from the outside. It carries a model of the machine, compares that model to what the machine is doing, and reasons through the delta. The error is the drift — the gap between what ought to be happening and what is. That reframing has stayed with me for 25 years.
This is Model-Based Programming — give the software a model of the physical system and let it reason about what to do, instead of hand-coding rules for every case you can think of. Brian Williams at MIT proved it in the hardest environment you could pick: NASA’s Deep Space One probe, where light-delay to Earth ruled out real-time human control. The spacecraft reasoned from a model of its own systems, diagnosed a stuck switch, and worked around it without anyone telling it how.
At the time, building the full model was quite complicated. We leaned on heuristics and simplifying assumptions to make it tractable. The idea was right — the economics weren’t there yet.
Today the concepts are wildly relevant, in the places doing what Williams did rather than borrowing the phrase “world model.” Wayve trains an explicit model of driving scenes — GAIA — that predicts how a scene plays out under different actions, and uses it to simulate situations too dangerous to drive into. DeepMind does something similar for virtual worlds: Genie 3 builds the world, and a separate agent, SIMA 2, plans and acts inside it. Yann LeCun’s JEPA work argues the same thing from first principles — intelligence needs a dedicated world-model module, kept apart from the thing that acts. Physical Intelligence and Figure get called “world-model” companies too, but their systems are closer to one big end-to-end network with nothing you can pull out and inspect. Same ambition as Williams, delivered as pattern-matching instead of modelling.
With the cost of modelling falling fast and the intelligence to build one skyrocketing, we’re all going to end up thinking like immobot designers — whether we use the word or not. Product design already works this way, if you squint: it’s the act of building a model of your users — their abilities, limits, goals — and shaping a system that optimizes against it. The part that gets skipped is the immobot part: checking the drift between that model and reality, continuously, in production, not just once at launch. MLOps has a name for this now — model or data drift — but it’s the same diagnostic problem Xerox and BMW were solving in 2002, aimed at a user model instead of a fuel sensor.
That’s the connection to AI software factories that matters most to me. A factory that’s only good at generating code is doing half the job. A key input has to be the model of the world its output is going to affect — call it a digital twin of the domain if you like — plus telemetry and diagnostics from the system in production. Synthetic users, generated from that domain model, let you pressure-test the output before a real person sees it. Without that loop, a software factory is a much faster way to generate code nobody’s checked against a model of what it’s for.
There’s a sharper version of this closer to home — Shelterwood is mine, and I’m the one running the factory I’m about to describe. In a recent post I borrowed a taxonomy from Antikythera’s Superdark Factory idea: Class 1 executes a plan I hand it, Class 2 takes a goal from me and plans its own way there (which is where my agents live today), and Class 3 — “Superdark” — would set its own objectives and move faster than I could describe. Model-based reasoning is a Class 2 thing, not Class 3. The Remote Agent didn’t invent its own mission; NASA set the goal and it planned around the parts nobody anticipated. That’s why I’ve stayed at Class 2 instead of reaching for the speed Class 3 promises.
I found out why the hard way. Three recent failures in my fleet of agents — one acting under my own credentials, a credential fragment that leaked into a transcript, one agent reaching permissions it had no business touching — never showed up as a bad outcome. I caught all three by checking actual state against a model of correct state: who should hold which key, whose name should be on an action. That’s drift, not a bad result, applied to a company instead of a machine — and it only worked because my agents are named. A Superdark factory erases that on purpose — agents are fungible, nothing carries an author — so nobody can slow it down by reviewing every decision. But Williams’s whole point was that you need named components in known states before you can reason about faults. Judging on outcomes alone is what let those three failures sit unnoticed. Naming my agents is what let me catch them as drift instead. That’s not autonomy solving the short-leash problem from the original article. It’s the same opacity, moved up a level, from the machine to the company.
25 years on, the lesson from a stuck switch on a spacecraft is the one I’d give anyone building with AI today: model first, generate second.