September 28, 2026 · 8 min read
A cheaper model per token is not a cheaper model per item
I spent a month testing a fast, cheap decision model as the core of a web-watching agent. It lost on every real task. The useful result was an architecture that barely needs a model at all.
- AI engineering
- Agents
- Research methodology
For the past month I have been testing Jev, TypeSafe's "System One" decision model, via OpenRouter as the engine for a persistent agent that watches websites and tells me what changed. Jev is not a chat model. You give it some state and a set of finite questions, such as "pick one of these options" or "give me a probability", and it answers in well under a second at a very low price per token. On paper that is exactly what an agent making thousands of small judgments needs, so I was excited to experiment with it.
The early work used small synthetic cellular-automata worlds that I built for the purpose. They were cheap to run and easy to control, but they couldn't tell me whether any of it would hold up on real websites. So I switched to live public pages and gave every experiment the same two rules.
The first rule was that every answer gets checked against ground truth. The pages I used were list pages: GitHub issue lists for the Python, Rust and VS Code projects, Hacker News, Lobsters, GitLab, Stack Overflow's newest questions, dev.to articles, Mozilla's Bugzilla, and a few others. Each entry on those pages (an issue, a story, a question or an article) has a title, usually an author, and often some tags or labels, and I'll call each entry an "item" from here on. All of these sites also publish the same lists through a public API, so I could check every answer against the site's own record of what was on the page.
The second rule was that Jev always went up against a strong general model on the same task. I used gpt-6-luna, a general-purpose LLM that costs $0.10 per million input tokens, which is cheap but not as cheap as Jev.
Reading the page
The first task was simply finding the items on a page. I loaded each page in a headless browser, split it into its visible pieces of text (every link, label, name and number the page shows), and asked for each piece's role. A piece could be an item's title, its author, one of its tags, or anything else, which on these pages means navigation links, buttons, timestamps, vote and comment counts, and section headers.
Jev's interface shaped how I could ask this. Each request carries some shared context plus up to sixteen separate multiple-choice questions, and Jev answers each question on its own. It can't read a whole page and hand back a list of items, so the natural fit was one question per piece of text ("what role does this piece play?"), sixteen pieces per request, and about thirty requests for a typical page of 450 pieces. gpt-6-luna has no such limit, so it got the whole page in a single call and wrote back the list of items directly.
Jev found nearly every title. The trouble was that it also called a lot of other things titles, and on the GitHub pages only four to six out of every ten pieces it labeled as a title actually were one. It kept picking up navigation links, page headings and pinned issues. gpt-6-luna scored 0.968, on a scale from 0 to 1 that penalizes both missed items and wrong ones (the F1 score, which I'll use for the rest of the post). My read on the gap is that every item on a list page repeats the same layout, and that repetition is the strongest clue about what is a title and what isn't. A model reading the whole page can see the repetition, while a model answering questions about sixteen pieces at a time mostly can't.
Writing the rules once
The most useful result came from a different idea. Instead of asking a model to read every page on every visit, I asked gpt-6-luna to look at one page once and write a small set of rules for that site, something like "the title is the main link inside each list row, and the author is the link next to it that points to a user profile". Ordinary code then applies those rules to every later visit of that site with no model calls at all.
On pages where each item is a list row or a table row, these rules matched gpt-6-luna reading the whole page. They failed at first on pages built from "cards", which is when each item is drawn as a stack of generic boxes with nothing in the page structure calling it a list row. Stack Overflow's question list and dev.to's article feed both work this way, and because nothing marked where one item ended and the next began, the rules had nothing to hang on. I fixed that by looking for groups of three or more sibling boxes with the same structure, which is what a repeated card looks like, and Stack Overflow's score went from 0.12 to 1.0. After that fix, the written rules tied gpt-6-luna on titles across all twelve pages in the final test, for roughly a quarter to a third of its cost on the first visit and nothing on later visits. For the six sites I captured more than once, rules written from the first capture still found every title on two later captures of the same pages.
Two ideas that sounded good didn't help. I tried letting gpt-6-luna revise its own rules by looking at what they extracted, without any answer key, and the revisions came out slightly worse because it had no reliable way to tell which version was better. I also tried Jev as a checker that scored each extracted item as plausible or not. Its scores were reasonable on average, but when I used them to throw out low-scoring items, it threw out every item from three of the sites, including Hacker News and crates.io.
Sorting items by topic
Jev's best chance looked like triage, which is a set of small yes-or-no questions. The items here were issues from the Python, Rust and VS Code GitHub projects, GitLab issues, Stack Overflow questions and dev.to articles, and each one already has labels that its maintainers or authors attached. Given only the title, the job was to predict those labels, for example whether a Python issue is labeled "interpreter-core", whether a Rust issue is labeled "C-bug", or whether a Stack Overflow question is tagged "python".
On a first set of 464 items Jev nearly tied, scoring 0.551 against gpt-6-luna's 0.560 at 63 percent of the cost, which looked promising. Since 464 items isn't much, I ran it again with 2,280 items from six sites, and this time I wrote down which labels to test before running either model so I couldn't be tempted to pick the ones that happened to work. For each site that meant up to eight of its most common labels, each needing at least fifteen examples, after excluding a written list of workflow labels. Workflow labels are things like "needs triage", which means no maintainer has reviewed the issue yet. Whether an issue carries that label depends on where it sits in the project's review process, and the title says nothing about that, so no model can predict it from the title alone.
At that size gpt-6-luna won, 0.614 to 0.565, and it came out ahead on five of the six sites. Jev cost $0.028 per thousand items against gpt-6-luna's $0.041, so it was 31 percent cheaper. I also tried giving Jev short descriptions of each label written by gpt-6-luna, which made its scores worse and its cost higher.
The small test had flattered Jev. I've learned the same thing in my trading research, where a result that looks good on a few hundred examples often doesn't survive a larger test.
Why the savings shrank
Jev's input tokens cost $0.042 per million, which is about 58 percent cheaper than gpt-6-luna's $0.10, so I expected a bigger gap. The difference comes from how much work each call does. gpt-6-luna labeled forty items in a single call. Jev's sixteen-question limit meant two items (with eight labels each) per request, and every request has to resend the shared instructions. By the time you count per item, which is what the bill actually tracks, Jev's price advantage was down to about 31 percent, and its answers were less accurate.
Jev is much faster per request. The median Jev request took 0.37 seconds, while a forty-item gpt-6-luna call took about 17 seconds. That speed matters when decisions arrive one at a time and each one needs an answer right away, for example routing an incoming support message to the right team the moment it arrives, or screening a comment before it gets posted. Watching websites isn't like that. My agent checks each page on a schedule and gets twenty to fifty items at once, so waiting 17 seconds for one batched answer is fine, and the batched answers were more accurate for a modest extra cost.
What I'm keeping
- The written rules. Paying a model once to write the rules for a site and then running them for free on every visit beat calling any model on every visit, however cheap the model was.
- A small database of what the agent has seen. It keys each item by its link, so it can tell what's new, what dropped off the page and what changed since the last visit. Across three captures it caught 75 percent of new items with no false alarms. Those captures were only about an hour apart, so I haven't yet seen how the rules and the database hold up when a site redesigns its pages, which is the next thing I want to test.
- Measuring cost per item. The price per token is what's on the pricing page, but the cost of each useful answer at the quality I need is what decides the design.
- Larger samples and rules written down in advance. Without the second, larger triage run I would have written up the near-tie as a win.
The whole series cost about fifty-six cents in model calls, which made it easy to rerun things at a larger scale whenever a result looked better than I expected.