🔍 Read the full analysis: 24 Ways To Use Jev When Designing AI Decision Models on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer published a Sept. 29 guide mapping 24 proposed uses for Jev, a tool that returns structured answers to narrow questions for software decision models. He says three uses are running in his publishing operation, 12 meet his fit test, seven need measurement and two are poor fits; the supplied source details only some of the examples.
Thorsten Meyer published a guide on Sept. 29 mapping 24 ways to use Jev in AI decision models, reporting that three applications are already live in his publishing operation. The piece matters to teams evaluating automated checks because it sets out conditions for deciding whether narrow, high-volume judgments are suitable for the tool, while Meyer’s performance and cost figures are his own reported results.
Meyer describes Jev as a system that accepts text or JSON state along with typed questions and returns structured answers for software to act on. The listed answer types are a yes probability, a choice with per-option probabilities and confidence, or a score on ordered levels with confidence. He says Jev does not write, summarize or extract information, and that a call containing the state and questions takes about 0.3 to 0.9 seconds. He gives a cost of about $0.04 per million input tokens.
The article reports that three checks are in use: assessing whether a story fits a site, checking whether an article is in English, and classifying a headline into one of 31 topics as a fallback when a primary large language model makes an error. Meyer says a scan of 78,889 articles cost $2.01; it identified 1,576 non-English items, of which 1,553 were fixed. For the topic classifier, he reports 89% agreement with a frontier language model overall and 97% to 99% agreement when Jev confidence was at least 0.8.
His proposed use cases include detecting whether a source has enough verifiable facts to support a report, matching products to a roundup, checking for affiliate or free-product disclosures, assessing headlines, and moderating comments. Meyer labels the source sufficiency check, product matching and headline review as needing measurement first. Disclosure checks and comment moderation are listed as strong fits. He calls same-event deduplication a poor fit after a canary test found no duplicates.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Fit
The guide’s central practical point is that a low-cost model call is useful only when it addresses a real, recurring decision. Meyer’s four-part fit test calls for high volume, a narrow question and low-cost errors, as well as evidence that the current heuristic is failing. This gives readers a way to distinguish a plausible application from a measured need.
Confidence-based routing is the proposed safeguard. Meyer says systems should act on clear answers and send uncertain cases to a more capable model or a person. In his example for story relevance, the gray zone continues through the existing publishing path rather than being blocked. That design can limit the effect of uncertain classifications, though the article’s reported agreement figures do not establish performance across other tasks or organizations.
The distinction between a promising use and a demonstrated one also matters for publishers and other operators. A missed disclosure may create compliance concerns, while an incorrect product match or headline flag may disrupt routine work. Meyer’s recommendations treat those risks differently, with human review for disclosure misses and headline scoring framed as a nudge rather than a sole publishing gate.
Meyer’s Test and Live Examples
Meyer recommends testing a proposed use against 300 to 500 real past decisions before integrating it into production. He says to compare results overall and by confidence band, then review 20 disagreements to determine which system was right. His suggested threshold is to wire in Jev only where the high-confidence band reaches 95% accuracy.
For rollout, the article advises putting the integration behind a separate feature flag that is off by default, trying it on 5% to 10% of units, and expanding only after a canary phase. These are Meyer’s proposed procedures, not reported findings from an independent evaluation. He also says a keyword rule should remain in place if it works, and that a use case should be revisited when measurement shows a problem.
The supplied article excerpt gives detailed tables for the three live publishing checks and the first several additional examples, but it ends during a section on commerce and customer operations. It does not provide the full set of 24 cases or supporting results for all of them. The headline count and overall fit categories are stated by Meyer; readers cannot assess every listed application from the available material.
“The superpower is confidence.”
— Thorsten Meyer, in the Sept. 29 article
Evidence Still Depends on Meyer
The article is a first-person account, and the supplied material contains no independent evaluation of Jev’s accuracy, speed or cost. Meyer does not provide the underlying test data, the definition of agreement with the frontier model, or a breakdown of errors beyond the confidence figures he reports for one 31-topic classification task. Agreement with another model also does not by itself establish that either answer was correct.
The excerpt does not include the full descriptions of all 24 proposed uses, nor does it name the 12 strong-fit cases and seven cases requiring measurement in a complete list. It says 15 uses are ready to build or already running, but the visible detail describes only three live cases and a partial set of additional proposals. The remaining applications, their evidence and the basis for categorizing them are not available in the supplied text.
Other operational details are also unspecified, including how the reported token cost was calculated, what data handling arrangements apply, and how the checks perform on material outside Meyer’s publishing workflow. The article does not establish whether the results will generalize to different content, question designs or error costs.
Measure Before Wider Rollout
Meyer’s recommended next step for teams considering a use is to replay several hundred past decisions, inspect disagreements and verify accuracy in the high-confidence group. He proposes enabling the tool behind a feature flag and starting with a small canary deployment after that review. These steps are guidance in the article; the supplied source does not announce a new product release or a rollout date.
For the seven cases Meyer says need measurement, the next milestone would be evidence that the existing rule or matcher fails often enough to justify replacement or supplementation. For the proposed deduplication check, his stated trigger is a measured duplicate problem. The article gives no schedule for publishing further results or details on the remaining use cases, so whether those evaluations or additions will follow is unclear.
Key Questions
What is Jev, according to Meyer?
Meyer describes Jev as a tool that takes text or JSON plus typed questions and returns structured answers, such as probabilities, choices or scores, that software can use to make decisions.
How many uses does Meyer say are ready to build or running?
He says 15 of the 24 mapped uses are ready to build or already running: three live, 12 strong fits, seven requiring measurement first and two poor fits. The supplied excerpt does not show the complete list.
What results does the article report from live use?
Meyer reports a scan of 78,889 articles for $2.01, identifying 1,576 non-English items and fixing 1,553. He also reports 89% overall agreement with a frontier model for a 31-topic classifier, rising to 97% to 99% at confidence of 0.8 or higher. These are his reported figures, not independently verified results.
How does Meyer recommend testing a proposed use?
He recommends replaying 300 to 500 past decisions, comparing outcomes by confidence band, reviewing 20 disagreements, and integrating the tool only where the high-confidence band reaches 95%. He then suggests a small canary rollout behind a feature flag.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
