AimieIQ is the natural-language interface to our Order-to-Cash platform. Ask it about your accounts, and it answers. Almost every question needs something different on screen, and there is no fixed page to design in advance, so we did the obvious thing and let the model design one each time.
We gave it a DSL (Domain Specific Language, or a small invented language for describing an interface), which the model writes, and our front end turns into components.
It worked. Users got charts, tables and contact cards assembled on the fly, for questions nobody had built a screen for.
We removed it anyway. The rules for writing that language ran to about 3,700 tokens, sat in front of the model on every turn, and were making the agent worse at the rest of its job. With the full rules loaded, it failed to carry a customer identifier through to the next step on four attempts out of four. In a matched test, taking the DSL away cut answer time by more than a third and got the model out of the business of writing numbers altogether.
The DSL we gave the model
The DSL was deliberately small. One component per line, each one named, and a single entry point our renderer followed to assemble the page. The catalog of components it could choose from was generated from our own design system, so the vocabulary the model saw always matched what we had actually built. Alongside it sat a set of instructions telling the model how to produce the right structure and a UI that made sense.
It is easy to see why people build it this way. A model that writes the interface can answer a question nobody designed for, without anyone building a template first. Streaming comes for free, because what the model writes is simply its reply, and the page assembles itself while the model is still writing.
Where it started to cost us
The DSL had to sit in front of the model on every turn, including the many with nothing to render. The token bill was the easy part. A long set of instructions competes with everything else you have asked the model to do.
We could measure that. At around 30,000 characters of instructions, the agent stopped carrying a resolved customer identifier through to the specialist that needed it, four attempts out of four. Halve the instructions and it carried the identifier every time. We saw the same pattern across three prompt sizes. The more room the DSL took, the worse everything else got.
The clearest example came from trying to fix a rendering fault by writing a clearer rule for it. The new wording added about a thousand characters. Across ten runs the fault it targeted became three times more common, and an unrelated check on whether the agent reached for the right tools fell by three quarters. Explaining the rule cost more than the fault did.
And when the model gets the syntax slightly wrong, nothing happens. That was the problem. A page with a missing entry point renders blank without raising an error, so the words the model wrote never reach the reader. A component defined but never attached is dropped with no error either, leaving a heading with nothing beneath it. A component whose arguments arrive named when the renderer wanted them ordered renders as empty space, which is how we once spent a while hunting for a table that had never been there. Each needed its own defensive code on the front end, written after we watched it happen to a real answer.
A page that renders nothing raises no error, so nobody finds out until a reader does.
Naming the fields that matter
What runs today inverts the arrangement. The specialist agent finishes with two things. First, a sentence or two of prose, and second, a list of the facts that belong on screen. It names those facts, and code renders them afterward.
A name here is what we call a reference, and it is nothing more than the tool the data came from and the field inside it, such as the overdue balance on an aging summary. The model picks which ones matter. It never writes the value.
There is no DSL to teach and nothing to parse. The only vocabulary the model needs is the field names it has already seen in the data it just read, so a name it could not have seen resolves to nothing instead of an invented figure.
Code does the rest. It looks up each named field in the data the turn already fetched, groups related facts into sections, picks a component for each, and formats every value by the type the data carries.
What the change bought
The most valuable property is one we did not set out for. The model no longer writes a number at all. Every figure is read from the source data as the page is built, so the prose and the tile beside it cannot disagree, and nothing can be mistyped on the way.
The page follows the question now. Ask for one figure and you get one figure; ask for a full account review and the same code assembles the full page, because the layout is composed from whatever the model named.
Turns got faster and lighter. On a matched pair of account summaries, wall-clock time fell by about 37% and the final generation went from around 3,900 tokens to 1,600. Treat those as indicative. They come from a single pair of traces, and our rule is two agreeing measurements before we believe a before and after.
Day to day, the work got smaller. Changing how something looks is a front-end change with no back-end release. Adding a figure is a line of configuration. A tool nobody has configured still renders something sensible, because the shape of the data is enough to pick a component. And a new capability no longer needs an output format, a template and a set of display rules. You tell the model which facts matter, and the rendering is solved.
About 37% faster on one matched pair of account summaries. Indicative only.
About 3,900 tokens down to 1,600.
One line of configuration, and the page picks it up.
Which side owns which decision
Relevance belongs to the model. It decides which facts answer the question and writes a sentence of prose to frame them. It never chooses a value, a component, or a format.
Presentation belongs to code, and so does the arithmetic behind any judgment. Risk flags are computed once from the underlying figures and frozen with the answer, so a record of what was flagged stays true after the thresholds move.
Two composition rules do more work than they seem to. A section appears only once the model has named enough of it, so a lone figure stays a lone figure while several become a section. Once it appears, it fills in its own required fields. Name a contact's email address and you get a card with their name on it. The rule completes a choice the model made and can never invent one.
flowchart LR
subgraph MODEL[Model]
direction TB
A["Prose answer"]
R["The facts that matter,<br/>named"]
end
subgraph SERVER[Server]
direction TB
P["Data the turn<br/>already fetched"] --> FL["Risk flags computed<br/>once and frozen"]
end
subgraph CLIENT[Rendering]
direction TB
RS["Look up each<br/>named fact"] --> SC["Group into sections,<br/>pick components"] --> BK["Page renders"]
end
MODEL --> CLIENT
SERVER --> CLIENT
classDef model fill:#ffe9e0,stroke:#ff5a28,color:#0C1B45
classDef plain fill:#ffffff,stroke:#BFC7D5,color:#8E97AE
class A,R model
class MODEL,SERVER,CLIENT plainThe step in between
Between the two there was a version we built and did not ship. The model filled in a typed object and our code turned it into the same DSL from a template. Moving layout into our hands was the right instinct, but it kept the DSL anyway, because the model still wrote it by hand whenever no template matched.
It also tied the page to the kind of request. In one trace a user asked for aging balances, the wrong deliverable was selected, the specialist still named exactly the three figures the user wanted, and the template discarded them in favor of a fourteen-field summary. That is a failure the current design cannot have, and it is why we went further.
The tradeoffs we accepted
Expressiveness is narrower, and that is the real price. A model that writes the interface can produce something for a question we never imagined. Composed rendering needs a way to present each shape of data it meets, and anything bespoke needs a new component.
The data became the contract. Every reference points at a field in an upstream system, so when one is renamed, the block that depended on it stops appearing. Tests cover the fields we rely on, so we find out quickly, which is not the same as preventing it.
Labels are ours now. The old approach had the model write them, so they arrived already translated, if inconsistently. Doing it properly means real translation keys, so adding a figure is a line of configuration plus a translation request. We think that is a fair price for labels we control.
The principle carries to other problems. Let the model decide what matters, and let code decide what it looks like. Working out which facts answer a question is the part that needs a model. The rest is layout, and layout we already know how to write.



-900x386.png)