🔍 Read the full analysis: A Practical Look At Jev: 24 AI Decision-Model Applications on ThorstenMeyerAI.com
TL;DR
Thorsten Meyer has mapped 24 possible uses for Jev, a tool that returns calibrated answers to typed questions so software can route routine decisions. He says three applications are running in his publishing operation, 12 meet his strong-fit criteria, seven need measurement and two are poor fits. The examples and performance figures come from Meyer; independent verification is not provided in the source.
Thorsten Meyer published a practical review of 24 applications for Jev, an AI decision tool, reporting that three are running in his publishing operation and classifying 12 more as strong fits. The article matters to teams considering automated decisions because it pairs each proposed use with a fit test and a rule for sending uncertain cases elsewhere.
Meyer describes Jev as a system that receives text or JSON plus typed questions and returns answers software can use directly. It does not, he says, write or summarize content. The available answer types include a yes-or-no probability, a choice with probabilities and confidence, and an ordered score with confidence. Meyer estimates a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
Three applications are live in Meyer’s publishing operation: checking story relevance to a site, detecting non-English articles, and using Jev as a fallback classifier across 31 topics. Meyer reports that one overnight scan covered 78,889 articles for $2.01, finding 1,576 non-English items and fixing 1,553. For the classifier, he reports 89% agreement with a frontier large language model overall, rising to 97%–99% when Jev’s confidence was at least 0.8. These are figures reported by Meyer; the source does not provide independent validation or a full evaluation methodology.
Among the publishing examples, Meyer marks disclosure checks and comment moderation as strong fits. He says a thin-source detector, product matching for roundups and headline-quality checks need measurement first. A same-event duplicate detector is labeled a poor fit because, according to Meyer, its canary test found no duplicates to address.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Help
The proposal targets work where people or larger models repeatedly make small, narrow judgments across many items. If the tool handles clear cases and sends uncertain ones to a person or a more capable system, teams may be able to apply checks more widely without making every decision depend on a costly review. Meyer’s examples include language screening, comment routing and content disclosure checks.
The limits are as relevant as the possible savings. Meyer’s own duplicate-detection example shows that a plausible use is not necessarily a real operational need: his test found no duplicates. His four-part test requires high volume, a narrow question, low-cost errors or a route for uncertain cases, and evidence that an existing heuristic fails. That final condition asks teams to show a measurable problem before adding another automated step.
Meyer’s Four-Part Fit Test
Meyer says confidence is central to how he expects Jev to be used. In a measurement on one 31-topic classification task, he reports 97%–99% agreement with a frontier LLM when confidence was 0.8 or higher, compared with 42% when confidence was below 0.5. The comparison is specific to that test, and the source does not describe its sample size or establish that the same accuracy applies to other tasks.
His proposed rollout starts with replaying 300 to 500 past decisions, comparing performance overall and by confidence band, then examining 20 disagreements to determine which system was right. He recommends wiring a use case only if the high-confidence band reaches 95%, putting it behind a feature flag that is off by default, and beginning a canary on 5%–10% of units. In this design, the application’s own code determines what action follows each answer.
„Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.“
— Thorsten Meyer
Evidence Still Needs Testing
The source is Meyer’s account of his own tool and deployments. It does not include an independent audit, full test data, or enough methodological detail to assess how representative the reported accuracy, cost and time figures are. Those figures should be read as Meyer’s measurements, not established results for every Jev use case.
The review also does not provide details for all 24 applications in the supplied material. It says seven uses need measurement because a failing heuristic has not been demonstrated, and two are poor fits, but the available source excerpt gives examples mainly from publishing. It remains unclear which applications outside that section are already deployed, what error rates they produce, and how users’ data is handled.
Measure Before Wider Rollout
Meyer’s next step for proposed applications is to test them against real past decisions, inspect disagreements and confirm that high-confidence results meet his stated threshold. He recommends starting any deployment behind a disabled-by-default feature flag and expanding through a small canary only after the test supports it. The article does not announce a product launch date or a broader release schedule.
For readers evaluating a similar system, the practical next milestone is evidence that an existing rule or workflow is failing at a meaningful rate. Until that is measured, Meyer’s framework treats an application as a candidate for testing rather than a demonstrated improvement.
Key Questions
What is Jev?
Meyer describes Jev as a tool that takes text or JSON and typed questions, then returns calibrated answers such as yes-or-no probabilities, category choices or ordered scores. Software can use those answers to route decisions.
How many of the 24 applications are already running?
Meyer says three are live in his publishing operation. He rates 12 as strong fits, says seven need measurement first and labels two poor fits.
What does Meyer say Jev costs and how fast does it respond?
Meyer estimates that one call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. These are figures from his article, not independently verified benchmarks.
How should teams decide whether to use Jev?
Meyer’s test looks for high volume, a narrow question, low-cost errors or a fallback for uncertain answers, and a visibly failing existing heuristic. He recommends replaying past decisions and checking accuracy by confidence band before deployment.
Source: ThorstenMeyerAI.com