G-CPU: We Removed the Vision Model From AI Agents. Entirely.
Enaptic is patent pending. U.S. Provisional Patent Application No. 64/140,784, filed August 25, 2026, covers five inventions built and proven inside a working multi-agent system: structural perception, active perception, sessionless execution, resident-context measurement, and supervised deployment control. This article introduces the first — and the idea that names this series: the G-CPU.
Born From a Stall
The system was not born from a theory. It was born from a stall. An agent in our fleet was blocked mid-task, reaching repeatedly for a vision tool that could not answer, and the way out was noticing that the answer had never been in the pixels to begin with — the drawing application already knew every layer’s name and bounds, in its own memory, one query away. The theory came after, to explain why the fix worked.
The GPU was re-deriving a fact sitting in another process’s memory.
Grown, Not Taken
Begin with a nine-line file: two partial differential equations, a feed-and-kill ramp, a grid size, a step count, twenty-four seeds placed by golden angle, a random seed, a palette. Run it anywhere numpy runs and the same organism appears, pixel for pixel — a Gray-Scott reaction-diffusion pattern, grown twice to hash-identical output. The genome is more the image than the pixels are. The pixels are phenotype; rasterization is development; and a vision model staring at the rendered output is reverse-engineering the genome from the organism. This is the founding epistemics of Enaptic Vision, and it yields a design test we apply everywhere: could you regrow the frame from the description? If yes, you have structure. If no, you have a summary.
What a Look Costs Today
A rendered screenshot submitted to a vision model costs three ways at once. It costs context: a single full-resolution capture consumes thousands of tokens of an agent’s bounded working memory, so an agent that verifies each action visually exhausts its window on perception alone and loses its working state to forced compaction. It costs truth: a model’s reading of pixels is probabilistic — it can report structure that is not there and miss structure that is, and verification built on testimony inherits its unreliability. And it costs the accelerator: every look is an inference, billed in latency and in the scarcest resource in the industry. These are not three small taxes. In sustained agentic operation, looking is the most frequent act the agent performs.
The Move Nobody Made
The efficiency literature has attacked the cost of inference with great sophistication — CPU offload, heterogeneous scheduling, split execution, co-packaged memory. All of it distributes the same forward pass across cheaper silicon. It makes the inference cheaper to run. None of it asks whether the inference needed to run. That is the move Enaptic Vision makes: not offloading the model, but removing it from the perception path — because the question a perceiving agent is actually asking is almost never a question about pixels. It is a question about structure, and structure has two honest sources, neither of which is a learned model: the application that produced the scene, and the arithmetic of the field itself.
Enaptic Vision: the Percept
The unit of perception is the percept: a plain-text structural frame with a fixed schema and a hard byte cap of roughly 1.5 kilobytes — about one percent of a small model’s context window. Perception, by construction, can never become a firehose. The frame is ordered the way biological vision is ordered: change first. The delta since the agent last looked leads; then a gist of the application and document; then a structural tree of elements with names, types, and bounds, trimmed to budget lowest-priority-first; then flags the daemon computes itself — occlusion, offscreen extent, unsaved changes; then the intent-diff, described below.
A background daemon — the optic nerve — polls the scene’s sources at about 2 Hz, hashes the state, and rewrites the current frame atomically whenever it changes, appending each delta to a bounded ring. The agent’s act of looking is therefore a file read completing in tens of milliseconds. “Instant” is an architecture decision, not a query speed. When an agent needs more than the gist, it asks: a focus request names one element and returns full detail for that element only. Detail is paid for when attended to, as in foveal vision — full state is never transmitted.
Structure is recovered by whichever of two paths applies, unified behind one schema. Source-first: where a producing application exists, adapters read its own object model — a layer tree, an accessibility tree, a document object model, a scene graph. The screen-reader precedent settles sufficiency: blind users operate full graphical applications on zero pixels. Structure-from-field: where no producer exists — a photograph, a scan — structure is computed from the pixel field by deterministic arithmetic: edges, connected components, character recognition, symmetry, centroid and luminance statistics. The consuming agent cannot tell which path produced its frame, and does not need to.
The element with no obvious precedent is the intent-diff: the perception of absence. A declared specification of intended content is kept beside the percept, and each cycle the daemon mechanically compares expectation against extracted structure and reports what should be present but is not. “The apple is missing a line” is perceived with no picture in the loop — because absence is only perceptible against a schema, and the schema is the declared intent. No rendered image, no reference image, and no model participates.
The whole apparatus wears model clothing. An endpoint exposes the percept store through the same API as a local model server: when the frame is fresh, it answers from structure and says so; when stale, it proxies honestly to a configured fallback; when neither, it reports unavailability rather than inventing. Nothing invented is a design principle, enforced in code. Agents consume structural perception without modification — they think they are talking to a model. In Enaptic’s own fleet, the fallback is unset, and no vision model is installed at all — perception runs on structure alone.
Evidence
Measurements from the working system; each figure below is recited in U.S. Provisional Patent Application No. 64/140,784.
- A look is bytes, not megabytes. The percept is hard-capped at ~1,500 bytes — about one percent of a small model’s context window — for a scene a screenshot renders in thousands of tokens.
- Aesthetic measurement without a model. Composition metrics — balance, centroid, thirds, focal convergence — computed CPU-only in tens of milliseconds.
- Structure survives occlusion. Occluded elements reported with bounds and an occlusion flag from source structure; no pixel visibility required.
- Structure, not summary — the regrow criterion. A generative description recovered from one frame: six parameters to 0.005–0.21%, hidden-state correlation 0.999934; run 1,000 steps past the observed instant, correlation 0.9997 against ground truth.
- One companion result belongs to the filing’s second invention, active perception, and is reported in its own paper: an agent stating, before acting, the exact post-action bounds of all twenty elements on a commercial application whose source it never saw — and scoring twenty of twenty.
The Economics
The cost structure of agentic perception has a shape: perception cost scales with how often an agent looks; judgment cost scales with how often it must decide; and in sustained operation an agent looks between every action and decides rarely. Under pixel-model perception, those two curves are fused — every look is priced as a decision. Enaptic Vision separates them. Looking becomes a sub-cent, sub-100-millisecond read that cannot exhaust the context window; deciding remains a model call, made only when something genuinely requires judgment. The gap between the fused curve and the separated one widens with session length and with agent count.
At the small end, the consequence is sovereignty: on constrained edge hardware, no vision model competes for memory, and every byte of accelerator capacity goes to the judgment model. At the large end, the consequence is bigger, and it is the argument of this paper’s title. In frontier-scale serving — many agents, perceiving continuously, each vision call drawing on the scarcest and most expensive hardware in the industry — relocating perception to general-purpose arithmetic removes a large and permanently recurring accelerator load. The benefit of the G-CPU move scales with the cost of the accelerator it relieves. The parties with the greatest economic incentive to practice structural perception are the operators of the largest GPU fleets. Graphics reasoning leaves the GPU because it was never a graphics problem.
Video falls out rather than being added. For interfaces, the delta stream is the video: a screen recording becomes an event log, natively legible to a language agent, roughly a thousandfold smaller than frame capture. And because percepts are bytes, an agent can watch a remote machine over a network at a cost streamed pixels could never meet.
The Boundary
What honestly remains for a learned model? Prediction of human response. Whether a composition will land as beautiful, trustworthy, premium — that is a fact about viewers, and it lives in model weights, not in the data. Everything else in the perception path is measurement. Our operating rule, enforced across our fleet: a pixel diff is evidence; a model’s read of pixels is testimony; never use testimony where evidence is available. Judgment is performed by a reasoning model over a structural description — a model that reads measurements, not one that stares at shadows.
Coda
Enaptic Vision runs today inside a working multi-agent system, where it was born from a stall and hardened by production. The mechanisms described here — with four companion inventions: active perception, sessionless execution, resident-context measurement, and supervised deployment control — are the subject of U.S. Provisional Patent Application No. 64/140,784, filed August 25, 2026.
G-CPU and Enaptic Vision are terms of art introduced by this paper.
Next in the series: Active Perception. One glance cannot tell you how a window behaves, but acting on it with known parameters can. An agent predicted where all twenty elements of a commercial application would land before it moved the window. It scored twenty of twenty. That story, and what it means for machines that learn interfaces the way a doctor reads a reflex, is next.