Jev, a model that never writes a word
How TypeSafe's Jev answers typed questions with probabilities instead of text, and why its output costs nothing.
Contents1 to 8
A model that does not write
TypeSafe AI released Jev on 15 September 2026, in early access.11TypeSafe, "Introducing System One Models & Jev", launch post by Diogo Almeida, 15 September 2026: "After two years in stealth"; definition; comparison table (LLM input "from $0.20 to $10 / MTok", output "~5x more expensive than input tokens"; Jev output "FREE (too cheap to meter)"; "70ms-500ms" end to end; "40x-200x faster"); RLCD; evals disclosure, including the "System One LLM wrapper", which "tends to be slower and more expensive than giving decisions without probabilities", and reference labels that bias answers "towards OpenAI and Anthropic's models"; hallucination chart note; FAQ on the names, on whether Jev is "just a smaller LLM", on public benchmarks and on training data. https://typesafe.ai/blog/introducing-system-one-models-and-jev The company calls it a System One model. On the same day TypeSafe ended two years in stealth and announced $40 million in seed funding, led by DCVC.1122DCVC, "TypeSafe emerges from stealth with a new way of doing AI", James Hardiman, 15 September 2026: "DCVC leads the $40 million Series Seed"; latency "less than 100 milliseconds". https://www.dcvc.com/news-insights/typesafe-emerges-from-stealth-with-a-new-way-of-doing-ai/
A language model answers by writing. Jev generates no text. The caller defines the possible answers in advance, and Jev returns typed values with probabilities attached. TypeSafe's own summary is "unstructured state in, typed probabilistic decisions out".11
The launch post's FAQ asks whether Jev is "just a smaller LLM", and answers that "Jev is neither small nor an LLM".11 On Hacker News, chief executive Diogo Almeida called it "technically not a language model (it doesn't generate language)".33Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026. https://news.ycombinator.com/item?id=49718437
The class name comes from Daniel Kahneman's Thinking, Fast and Slow, published in 2011.44Daniel Kahneman, Thinking, Fast and Slow, Penguin Random House listing (published 25 October 2011), the page TypeSafe links. https://www.penguinrandomhouse.com/books/89308/thinking-fast-and-slow-by-daniel-kahneman/ TypeSafe borrows the book's split between "fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning".11 The launch post concedes that the label has also implied error-prone thinking, and holds that System One models "can be made more reliable than its alternatives".11
Jev is named after William Stanley Jevons. TypeSafe expects machine intelligence to follow the path of coal, whose use grew as steam engines became more efficient.11 The press release calls the name "a nod to Jevons Paradox".55TypeSafe press release, Business Wire, 15 September 2026, read in the verbatim copies on Morningstar and Yahoo Finance: "Founded in 2024 and headquartered in San Francisco"; "a nod to Jevons Paradox"; "less than 100 milliseconds of latency"; "up to 100 times faster and less expensive". https://www.morningstar.com/news/business-wire/20260915525333/typesafe-ai-emerges-from-stealth-with-40m-in-funding-with-new-model-for-composable-ai and https://finance.yahoo.com/technology/ai/articles/typesafe-ai-emerges-stealth-40m-190000776.html
TypeSafe was founded in 2024 and is based in San Francisco.55 Almeida is its chief executive. The chief operating officer, Sasha Sheng, was a research engineer at Meta and its FAIR lab, and the chief technology officer, Erik Gafni, has founded companies before.66TypeSafe team page, undated, read on 3 October 2026. https://typesafe.ai/team
Almeida is a co-author of InstructGPT, the 2022 paper on training language models to follow instructions with human feedback, and the fourth of its 20 authors.77Ouyang et al. (2022), "Training language models to follow instructions with human feedback", arXiv 2203.02155, 4 March 2022. Almeida is the 4th of 20 authors. https://arxiv.org/abs/2203.02155 TypeSafe's team page says he "co-invented RLHF and InstructGPT".66 He is not among the authors of the 2017 paper by Christiano and colleagues that introduced deep reinforcement learning from human preferences.88Christiano et al. (2017), "Deep reinforcement learning from human preferences", arXiv 1706.03741, 12 June 2017. Almeida is not among the six authors. https://arxiv.org/abs/1706.03741
In the launch post, Almeida writes that at OpenAI he "helped build the methods" that made language models useful at following instructions. That work "ended up as the research behind ChatGPT".11
Three kinds of question
Jev answers at one endpoint, POST /v1/systemone.99TypeSafe docs, API reference, read on 3 October 2026: one endpoint; up to 255 options per Choice, each with an optional description; at least two and up to 10 levels per Score; the Choice example on the payouts ticket (0.88, 0.12, 0.0, confidence 0.81, 318 input tokens) and the Score and Noul examples on the same ticket; the four example responses on the page report 18 to 34 output tokens. https://docs.typesafe.ai/api A request carries its input in a field called state, which can be a string, a JSON object or an array. A second field names the model version, such as jev-latest, and a third, questions, maps names chosen by the caller to typed questions.99 The names are for the caller's code and "are not sent to the model".1010TypeSafe docs, Primitives, read on 3 October 2026: question IDs "are not sent to the model"; "One question's answer is not hidden context for another". https://docs.typesafe.ai/primitives
There are three question types, and one request can mix them.99
A Choice picks one option from a list the caller writes. Each option can carry a short description, and a Choice allows up to 255 options.99 The answer names the winning option and gives a probability for every option. The probabilities sum to 1.99
A Score places the answer on an ordered scale of 2 to 10 levels.99 Its value is a probability-weighted average of the level numbers, so it can land between two levels.
A Noul is a yes or no question. The name is short for Bernoulli.1111Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026: "noul" is short for bernoulli and maps to if-statements, choice maps to a match statement and score maps to sorting. https://news.ycombinator.com/item?id=49718407 Its answer is one number, the probability of yes.99
On Hacker News, Almeida mapped each type to ordinary code. A Noul maps to an if-statement and a Choice to a match statement, while a Score maps to sorting.1111
The documentation's example is a support ticket whose whole state is one line, "Help! My payouts have been failing for 3 days." The question reads "Which team should handle this?", and the options are billing, technical and sales.99
In the documented response Jev picks billing, with 0.88 on billing, 0.12 on technical and 0.0 on sales. It reports a confidence of 0.81, and the request used 318 input tokens.99
Confidence is computed from the probabilities, and TypeSafe publishes the formulas.1212TypeSafe docs, Confidence, read on 3 October 2026. Choice: (p_max - 1/n) / (1 - 1/n). Score: 1 minus the probability-weighted distance from the most likely level, divided by the same measure for a uniform spread, floored at 0. Noul: the docs suggest |2p - 1|. https://docs.typesafe.ai/confidence For a Choice with n options, it is the top probability minus 1/n, divided by 1 minus 1/n. It is 1 when one option holds all the probability and 0 when the probability is spread evenly.1212 Applied to the ticket's rounded probabilities, the formula gives 0.82, one hundredth above the printed 0.81.1313Arithmetic on the docs examples. Choice: (0.88 - 0.333) / (1 - 0.333) = 0.547 / 0.667 = 0.82, printed as 0.81. Score: 0 x 0.0 + 1 x 0.95 + 2 x 0.05 = 1.05; confidence 1 - 0.05 / 0.667 = 0.925, printed as 0.92. https://docs.typesafe.ai/api and https://docs.typesafe.ai/confidence
A Score's confidence measures how far the probability spreads from the most likely level, against a uniform spread.1212 The docs' Score example asks "How frustrated is the customer?" about the same ticket, on three levels from Calm to Very angry. It returns 1.05, just above the middle level, with a confidence of 0.92.1313 A Noul carries no confidence field. Where an equivalent is needed, the docs suggest the absolute value of 2p minus 1.1212
The docs say every question in a request "is evaluated in parallel and in isolation against the same state in one go", and that adding questions "barely changes the response time".1414TypeSafe docs, Introduction, read on 3 October 2026: questions evaluated "in parallel and in isolation"; adding questions "barely changes the response time" and "does not create context-rot". https://docs.typesafe.ai/introduction One question's answer "is not hidden context for another".1010
- state
Help! My payouts have been failing for 3 days.
- questions
- departmentchoiceWhich team should handle this?criteria
- billing
- Payments, invoicing, refunds
- technical
- Bugs, outages, integrations
- sales
- Pricing, upgrades, new accounts
- choice
- billing
- probabilities
- confidence
- 0.81
- CONFIDENCE
- (pmax − 1/n) / (1 − 1/n)
- (0.88 − 1/3) / (1 − 1/3) = 0.820
Writing and deciding
In TypeSafe's launch table, an LLM "Generates one token at a time, each conditioned on the last".11 Its output is a string, and strings "can be anything", hallucinations and refusals included. Software that needs a decision must parse and validate the string before it can use it.11
OpenAI launched Structured Outputs on 6 August 2024. Its changelog says model outputs "now reliably adhere to developer supplied JSON Schemas".1515OpenAI API changelog, entry of 6 August 2024: "Launched Structured Outputs", read on 3 October 2026. https://developers.openai.com/api/docs/changelog Simon Willison wrote the same day that OpenAI was presumably using a known technique, which works on next-token selection so that only tokens valid under the schema can be chosen.1616Simon Willison, post on OpenAI's Structured Outputs, 6 August 2024. He writes that OpenAI is "presumably" using the same trick as jsonformer and llama.cpp grammars, which interact with next-token selection so that only tokens matching the schema are chosen. https://simonwillison.net/2024/Aug/6/openai-structured-outputs/ The model still writes its answer one token at a time.
Almeida argues that the masking makes models worse. On Hacker News he wrote that constrained decoding makes them "dumber", because a model that puts any probability on an invalid token is "by definition confused".1717Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026: constrained decoding makes models "dumber", because "simply masking logits is insufficient": a model that assigns probability to an invalid token "is by definition confused". https://news.ycombinator.com/item?id=49718849 The home page FAQ says that forcing an LLM into a format "can leave some of its intelligence on the table".1818TypeSafe home page and FAQ, undated, read on 3 October 2026. https://typesafe.ai/
Jev has no sequential output at all. Almeida wrote that strings and "all sequential data structures" are not allowed, so that "all outputs can be computed in parallel (thus no output token cost)".1919Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026. https://news.ycombinator.com/item?id=49719122 In the launch table, Jev "Generates all outputs in a single query".11
The difference grows with the number of questions. An LLM asked thirteen questions writes all thirteen answers into one JSON object, token after token. Each of those tokens is billed.11 Jev returns the thirteen answers from one call, and output is priced at zero.11
TypeSafe's parallel-questions cookbook asked 13 questions about the Wikipedia article on the GDPR, once in a single call and once as 13 separate calls to Jev. The single call came out "12.2x cheaper and 10.0x faster", with no change in the answers. The saving comes from sending the document once, in one round trip, instead of thirteen times.2020TypeSafe docs, Parallel questions cookbook, read on 3 October 2026: 13 questions over the GDPR article, asked in one call and as 13 single-question calls, "12.2x cheaper and 10.0x faster with no change in answers"; "The document dominates every request", so 13 calls pay for it 13 times, in 13 round trips. The Primitives page describes the same cookbook as 11.5x cheaper and 9.6x faster. https://docs.typesafe.ai/cookbooks/parallel_questions and https://docs.typesafe.ai/primitives
STATEHelp! My payouts have been failing for 3 days.
13 questions. The LLM writes 73 tokens, one per step, so 73 steps. Jev returns all 13 answers in 1 step.
SCHEMATIC. STEPS, NOT MILLISECONDS.
ONE STEP: ONE TOKEN FOR THE LLM, ONE QUERY FOR JEV. THE ANSWERS ARE MADE UP. UP TO 13 QUESTIONS, AS IN TYPESAFE’S PARALLEL-QUESTIONS COOKBOOK.
TypeSafe credits a new model architecture and a parallel sampler, along with a training method it calls Reinforcement Learning for Calibrated Decisions, or RLCD.11 Its docs present RLCD as a third kind of post-training after RLHF and RLVR, one that "trains TypeSafe to return decisions and calibrated probabilities instead of generated text".2121TypeSafe docs, Machine learning primer, read on 3 October 2026: RLCD after RLHF and RLVR; calibration as outcomes given 0.2 happening about 20% of the time. https://docs.typesafe.ai/introduction/machine-learning-primer
TypeSafe has published no paper on RLCD, and neither the model's size nor its base model.2222No RLCD paper, method description, model size, base model, revenue or cost figure appears on TypeSafe's site, blog or docs as read on 3 October 2026. https://docs.typesafe.ai/llms-full.txt On Hacker News, Almeida wrote that the architecture "is close to the chest for now", and that the team had talked about writing a paper.2323Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026: "architecture is close to the chest for now, but we have talked about writing a paper". https://news.ycombinator.com/item?id=49718824
TypeSafe says it makes all of its training data itself and does not train on customer data. The same weights serve every account.112424TypeSafe docs, Models, read on 3 October 2026: price $42 per Btok and $0.042 per Mtok, "Charged per input token. Output tokens are free."; rate limits; "64k tokens per request; 32k tokens for state plus the longest question"; text-only input; languages; current model jev-1.13.0 with both aliases; not fine-tuned on customer data, the same weights for every account. https://docs.typesafe.ai/models When the Latent Space podcast put it to Almeida that all of the data is synthetic, he answered "Yep".2525Latent Space, interview with Diogo Almeida, published 21 September 2026, with transcript: synthetic data at 00:22:15; a trillion tokens a day, "even at night", at about 00:35:32; versions, and a possible temporary long-term support for Jev 1.13.0, at about 00:48:11. https://www.latent.space/p/jev
The bill
Jev costs $42 per billion input tokens, or $0.042 per million, and "Output tokens are free."2424 The launch post lists output as "FREE (too cheap to meter)".11 The API still counts output tokens. The example responses in the API reference report between 18 and 34 of them, and none is billed.992424
The launch post gives LLM input prices "from $0.20 to $10 / MTok", with output at about five times the input price.11 Against that range, Jev's input is between 4.8 and 238 times cheaper.2626Arithmetic on TypeSafe's prices: $0.20 / $0.042 = 4.76; $10 / $0.042 = 238.1; 238 x $0.042 = $9.996. https://typesafe.ai/blog/introducing-system-one-models-and-jev and https://typesafe.ai/ The home page quotes the top end, "238x Lower input price than Claude Fable 5.1", which puts Fable 5.1 at about $10 per million input tokens.18182626
The comparison covers input only. An LLM answering in JSON is also billed for every token it writes, at about five times its input rate in TypeSafe's table.11
The docs' ticket used 318 input tokens.99 A million such tickets would cost $13.36 in input at Jev's price. At the two ends of the LLM range the same input would cost $63.60 and $3,180, before any output.2727Arithmetic at list prices, input only: 318 tokens x 1,000,000 tickets = 318 million tokens; 318 x $0.042 = $13.36; 318 x $0.20 = $63.60; 318 x $10 = $3,180. https://docs.typesafe.ai/api and https://typesafe.ai/blog/introducing-system-one-models-and-jev
The models page sets a rate limit of "100K tokens per second / 40 requests per second". The limits "can change without notice", and higher ones are available on custom and enterprise plans.2424 At the token cap, one account could send 8.64 billion tokens a day, which would bill $362.88.2828Arithmetic on the published limits: 100,000 tokens x 86,400 seconds = 8.64 billion tokens a day; 8.64 x $42 = $362.88; 40 requests x 86,400 = 3,456,000 requests a day; 40 x 318 tokens = 12,720 tokens a second, under the 100,000 token cap. https://docs.typesafe.ai/models At the request cap it could send 3,456,000 requests a day. For tickets the size of the docs example the request cap binds first, at 12,720 tokens a second.2828 Each request can carry many questions, so the number of decisions can be higher.
A request may hold "64k tokens per request; 32k tokens for state plus the longest question".2424 OpenRouter's listing gives a single "32,000 token context window". Cloudflare, which released a rival model on 1 October, contrasts its own 64k window with "Jev's 32k".2929OpenRouter, Jev Latest listing, read on 3 October 2026 ("32,000 token context window"); Cloudflare, Clef announcement, 1 October 2026 ("our model has a 64k context window (compared to Jev's 32k)"). https://openrouter.ai/~typesafe/jev-latest/ and https://blog.cloudflare.com/clef-decision-models/ Both figures match the narrower of TypeSafe's two budgets.2424
Input is text only. English is the main training language, and other languages, CJK scripts included, are handled "but not equally well".2424
TypeSafe has given two answers on whether the price covers its costs. The launch post says TypeSafe "can't prove it isn't subsidized", and expects the price "to go down, not up".11 The home page FAQ, read on 3 October, says TypeSafe can "serve Jev profitably" at current prices.1818 TypeSafe has published no revenue or cost figures.2222
Default rate limit at this state size: 3.46 million decisions a day.
Set by 40 requests a second
Above the default rate limit; higher limits on custom plans.
1,000 tokens a decision, 1,000,000 decisions a day. Jev $42 a day, $1,260 over 30 days. LLMs $200 to $10,000 a day, $6,000 to $300,000 over 30 days.
List prices for input tokens only. Output is left out: Jev’s is free, while LLMs bill it too, at about five times the input price in TypeSafe’s launch table. One request per decision; default limits per account: 100K tokens and 40 requests a second.
Fast and cheap, and how accurate
The launch post says TypeSafe chose not to publish results on public benchmarks.11 Its own evidence is a set of workflow evals at evals.typesafe.ai, with four workflows that run from security incidents to customer service.3030TypeSafe, Workflow evals, read on 3 October 2026. Each point averages a setup's accuracy, cost and time per item over four workflows with equal weight. The chart prints short labels without vendor or version: Jev 67.8%, $0.0004, 0.4 s; sol 74.1%, $0.0836, 23.3 s; opus 5 73.1%, $0.1761, 37.8 s; terra 67.9%, $0.0304, 10.1 s; sonnet 5 67.8%, $0.1174, 78.1 s; luna 66.8%, $0.0033, 12.9 s; DS v4 pro 65.5%, $0.0413, 86.5 s; DS v4 flash 64.4%, $0.0059, 51.9 s; haiku 4.5 53.6%, $0.0195, 12.5 s. By workflow (Security Incidents, Agent Trace Observability, Invoice Processing, Customer Service), Jev scores 61.7, 71.6, 61.8 and 76.0, sol 62.5, 76.6, 79.1 and 78.3. The chart also plots each LLM run from a standalone prompt; those runs score lower and are left out here. https://evals.typesafe.ai/ The LLMs answer through TypeSafe's System One LLM wrapper, which makes them return decisions in the format of Jev's API. TypeSafe calls it the most accurate way to get decisions from an LLM, and says it tends to be slower and more expensive than decisions without probabilities.11
The answer key comes from other models. The reference labels are "an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking", and every other model runs at its provider's default reasoning setting.3030 The launch post adds that the workflows were built by TypeSafe's model-capabilities team, "so some bias could exist". It also says the choice of reference labels favors OpenAI's and Anthropic's models.11
Averaged over the four workflows, Jev scores 67.8%, at $0.0004 and 0.4 seconds per item.3030 The best mean belongs to a setup the chart labels sol, at 74.1%. Opus 5 follows at 73.1%, and terra scores 67.9%. Sonnet 5 ties Jev at 67.8%, at $0.1174 and 78.1 seconds per item, and Haiku 4.5 comes last at 53.6%.3030
The gap depends on the workflow. On customer service Jev scores 76.0% and sol 78.3%. On invoice processing Jev scores 61.8% and sol 79.1%.3030
Jev's advantage is cost and time. Plotted against either, Jev sits on the Pareto frontier with terra and sol.3131Arithmetic on the eval values. Pareto frontier, by cost and by time alike: Jev (67.8%), terra (67.9%), sol (74.1%); every other setup is beaten on both counts by one of these. Terra against Jev: 0.0304 / 0.0004 = 76; 10.1 / 0.4 = 25.3. Cost: Opus 5, 0.1761 / 0.0004 = 440. Time: Sonnet 5, 78.1 / 0.4 = 195; Opus 5, 37.8 / 0.4 = 94.5. Means of the eight LLM setups: cost 0.4775 / 8 = 0.0597, and 0.0597 / 0.0004 = 149.2; time 313.1 / 8 = 39.1, and 39.1 / 0.4 = 97.8. Jev's $0.0004 is rounded, so the cost ratios are approximate. https://evals.typesafe.ai/ Terra beats Jev by 0.1 points at 76 times the cost and 25 times the time per item.3131
Means over TypeSafe’s four workflows, all nine setups run as workflows. Accuracy is agreement with reference labels made by averaging GPT-6 Astra and Claude Fable 5.1 at high thinking. The workflows were built by TypeSafe. Source: evals.typesafe.ai, fetched 3 Oct 2026.
The home page headline reads "193.6x Faster, 444.6x Cheaper", based on "workflows for System One tasks".1818 The launch post says the figures come from these evals and sit "on the higher end of real world gains".11 Neither page names the model behind the comparison.
The two numbers fit two different models. Opus 5 costs about 440 times as much as Jev per item, close to 444.6.3131 Sonnet 5 takes 195 times as long, close to 193.6, while Opus 5 takes only 94.5 times as long.3131 Against the average of all eight LLM setups, Jev is 149.2 times cheaper and 97.8 times faster.3131
TypeSafe's sources also disagree on speed. The launch post gives an end-to-end time of "70ms-500ms" and says its published evals are run from laptops on the West Coast. Elsewhere the same post speaks of "100ms speeds" for real-time apps.11 The docs' use-case map says "real-time speeds (150ms)".3232TypeSafe docs, Use-case map, read on 3 October 2026. https://docs.typesafe.ai/concepts/use-case-map The press release says "less than 100 milliseconds of latency", and DCVC repeats it.5522 The press release also says "up to 100 times faster and less expensive", where the launch post says "40x-200x faster".5511
Some users published their own figures. A developer posting as @nutlope summarized 1,018 papers with DeepSeek V4 Flash and classified them with Jev, for "$0.08 total cost and 256ms median end-to-end latency".3333Post by Hassan (@nutlope) on X, 17 September 2026, read as quoted in Diogo Almeida's post. Each paper was summarized with DeepSeek V4 Flash before its title, summary and 24 possible topics went to Jev. https://x.com/CompleteSkeptic/status/2100443429281120508 Guillermo Rauch, Vercel's chief executive, wrote that in Vercel's fx safety reviewer Jev is "up to 18x faster (p95)" and more accurate than the GPT Luna model the reviewer then ran on.3434Guillermo Rauch, post on X, 16 September 2026: the fx reviewer "runs on GPT Luna today". https://x.com/rauchg/status/2100307962262872105
A preprint by Ibrahim and Zaki, not yet peer reviewed, tested Jev on social-science annotation tasks. Jev trailed the best LLM for each task on 14 of 15 tasks, by a median 11.6 points of macro-F1, a score that weighs every class equally.3535Ibrahim and Zaki, "Evaluating Decision Models for Text Annotation in Computational Social Science", arXiv 2609.24574, 21 September 2026 (v1), preprint. The abstract calls Jev "the first commercial decision model"; the full text names it. Figures from the abstract and the calibration table: Jev's median expected calibration error 0.157, against 0.066 for Opus 5, 0.073 for Fable 5.1 and 0.075 for Sonnet 5, whose confidence is verbalized; the abstract prints the pair as "0.157 against 0.066". Median accuracy 0.815 at or above 0.9 confidence; routing low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost". https://arxiv.org/abs/2609.24574 and https://arxiv.org/html/2609.24574 Its measured cost was a median 44 times lower.3535
Calibration
A model is calibrated when its probabilities match how often things happen. TypeSafe's docs give the standard test. Outcomes given a probability of 0.2 should happen about 20% of the time.2121 Calibration is a property of many answers together. It "does not guarantee that an individual answer is correct".3636TypeSafe docs, System One, read on 3 October 2026. https://docs.typesafe.ai/concepts/system-one
TypeSafe's launch table explains why it trains for calibration. "If a model can do a task 95% of the time but doesn't say when it's in the 5%, it can't automate that task."11
The docs turn that into routing by confidence. Code acts alone on confident answers and sends the rest to a person or another system.1212
Geifman and El-Yaniv brought the same idea to deep networks in 2017, under the name selective classification. Their classifier rejects the inputs it is unsure about, so as to meet a level of risk set by the user.3737Geifman and El-Yaniv (2017), "Selective Classification for Deep Neural Networks", arXiv 1705.08500, 23 May 2017: the classifier "rejects instances as needed, to grant the desired risk". https://arxiv.org/abs/1705.08500
Language models can be calibrated in the right setting. Kadavath and colleagues found in 2022 that larger models "are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format".3838Kadavath et al. (2022), "Language Models (Mostly) Know What They Know", arXiv 2207.05221, 11 July 2022. https://arxiv.org/abs/2207.05221 A Choice is a multiple-choice question, and a Noul a true-or-false one.
Neural networks are often miscalibrated out of the box. Guo and colleagues reported in 2017 that modern networks "are poorly calibrated", unlike those of a decade before. A simple fix, temperature scaling, was "surprisingly effective".3939Guo, Pleiss, Sun and Weinberger (2017), "On Calibration of Modern Neural Networks", arXiv 1706.04599, 14 June 2017. Temperature scaling is described as "a single-parameter variant of Platt Scaling". https://arxiv.org/abs/1706.04599
TypeSafe publishes no calibration measurement for Jev. Neither its docs nor its launch post give an expected calibration error or a reliability diagram.4040No expected calibration error, reliability diagram, Brier score or log-loss for Jev in TypeSafe's full docs or launch post, searched on 3 October 2026. https://docs.typesafe.ai/llms-full.txt The only published comparison of confidence and accuracy is a cookbook on SEC filings, where a cut at 0.9 confidence splits 60 filings in half. The 30 confident answers were right 27 times, and the other 30 were right 12 times.4141TypeSafe docs, Classification using confidence cookbook, read on 3 October 2026: 60 filings, a cutoff of 0.9, the confident half right "90% of the time", the other half "40%", so 27 of 30 and 12 of 30; "Numbers below came from jev-1.12 on 2026-08-12". https://docs.typesafe.ai/cookbooks/classification_using_confidence Those runs date from August and used jev-1.12. The model now serving is 1.13.0.41412424
The outside measurements so far are preprints. Ibrahim and Zaki found Jev's confidence better calibrated than the confidence stated in words by 16 of 19 LLMs. Three frontier Claude models did better. The best of them, Opus 5, had a median expected calibration error of 0.066, against 0.157 for Jev.3535 Items at or above 0.9 confidence had a median accuracy of 0.815. On one task, empathy in peer-support dialogues, Jev reported high confidence at near-chance accuracy.3535
Ibrahim and Zaki also tested the routing the docs describe. In their abstract, sending low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost".3535
Rafe and Das, in a second preprint, used Jev to code 195,857 Texas police crash narratives with a 27-question schema. Against human labels it reached an F1 of 0.908. Of two frontier LLMs run on the same records, one scored 0.059 higher and the other was indistinguishable from Jev.4242Rafe and Das, "Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)", arXiv 2609.24052, 21 September 2026, preprint. Figures from the abstract: F1 0.908; "One frontier model gains 0.059 and the other is indistinguishable from it"; "Recalibration on the same labels reduces calibration error by a factor of 3.3". The full text fits the recalibration map to Jev's probabilities. https://arxiv.org/abs/2609.24052 Recalibrating Jev against the same human labels cut its calibration error by a factor of 3.3.4242
OVERCONFIDENT MODEL, THRESHOLD 0.90
- ANSWERED AUTOMATICALLY
- 61.8%618 OF 1,000
- WRONG AMONG THEM
- 9.9%ITS OWN SCORES PROMISE 2.2%
- SENT ON
- 38.2%TO A PERSON OR A LARGER MODEL
SYNTHETIC DATA, MADE TO ILLUSTRATE. NOT JEV’S NUMBERS.
Overconfident model, threshold 0.90. 618 of 1,000 decisions answered automatically, 9.9% of them wrong. Its own scores promise 2.2%. 382 sent on.
Where it breaks
TypeSafe's docs keep a page on where jev-1.13 fails, last reviewed on 17 September. It lists nine failure modes:4343TypeSafe docs, Model jaggedness, jev-1.13, "Last reviewed 2026-09-17", read on 3 October 2026. https://docs.typesafe.ai/model-jaggedness/jev-1.13
- literal reading;
- math and numbers, counting included;
- comparing dates and times;
- indirection;
- a large state full of irrelevant detail;
- adversarial content;
- contradictory instructions and criteria;
- common-sense structural invariants;
- generation.
On adversarial content the page says that "State is data, and jev-1.13 does not treat it as hostile by default". An injected instruction "can move the answer", and TypeSafe expects to improve on this.4343
On large inputs it says "Jev suffers from context rot, so unrelated material in the state costs you accuracy".4343 The docs' introduction says that adding questions "does not create context-rot".1414 The first sentence concerns the state, and the second concerns questions, which are evaluated in isolation.1414
The same page shows structural invariants that do not hold. One ticket reads "I was charged twice for the same order. Can someone look into this?" Asked as a Noul whether the customer is asking for a refund, Jev gave 0.72. Asked whether the customer is asking for "something other than a refund", it gave 0.47.4343 Read as one probability and its complement, the two would sum to 1.00. They sum to 1.19.4343
On another ticket, "I'm not happy with the fit. What are my options here?", the refund question was asked twice. As a Noul it got 0.22, and as a Choice between yes and no it got 0.01 for yes.4343
The names of the options also move the answer. In a preprint titled "Type-Safe Is Not Error-Free", Sun and Xu changed only which option name went with which description. On the hosted Jev, the swap moved the AUC from .8146 to .5806, where 0.5 is chance, and caused "24x as many answer flips as its test-retest floor".4444Sun and Xu, "Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It", arXiv 2609.26758, 22 September 2026 (v1), preprint, figures for the hosted model from the abstract. https://arxiv.org/abs/2609.26758 The type-error rate stayed at 0%.4444
The launch post says Jev "can't hallucinate". Its hallucination chart notes that TypeSafe's own figure "is not empirical", because "Schema matching is guaranteed".11 The claim concerns the form of the answer. A Choice always returns one of the options it was given, and a Noul always returns a probability.99
The home page FAQ says the same. Asked "Can Jev still get things wrong?", it answers "Yes. Jev guarantees the shape of its answers, not that every decision is correct."1818
When a Hacker News commenter called the claim wrong, Almeida replied that a wrong but valid value "is likely true of all ML". In another reply he asked whether a linear classifier hallucinates.4545Diogo Almeida, comments in the Hacker News launch thread, 15 September 2026. https://news.ycombinator.com/item?id=49718767 and https://news.ycombinator.com/item?id=49719080
The FAQ names tasks that may suit large reasoning models better, such as "complex mathematics or chess-like planning".1818 The docs say Jev is "not a drop-in replacement" for the LLM behind coding agents.4646TypeSafe docs, Coding agents, read on 3 October 2026. https://docs.typesafe.ai/introduction/coding-agents
The first two weeks
In the Latent Space interview published on 21 September, Almeida said TypeSafe had passed a trillion tokens a day, with machines calling the model "even at night".2525 At list price, a trillion input tokens would bill $42,000. TypeSafe has not said how much of that traffic is paid, and it has published no revenue.4747Arithmetic, illustrative only: 10^12 tokens x $42 per 10^9 tokens = $42,000 a day, if every token were billed at list price. The date of the milestone and the share of paid traffic are unknown. https://www.latent.space/p/jev and https://docs.typesafe.ai/models
- 15 SEPLaunch
- 17 SEPOff the waitlistAlmeida posts that TypeSafe has "jev-ed 140k off the waitlist" in under 36 hours.49
- 18 SEPVercelVercel says Jev "was adopted faster than any other model in AI Gateway history", reaching about 13% of teams on its first day.50
- 20 SEPNo waitlist"Jev is now available to everyone. No waitlist."51
- 22 SEPSignups pausedNew signups are paused after an "immense swell of demand". Existing accounts keep working.52
- 27 SEPSignups reopenAnyone can sign up again. Free credits are switched off for new users because of "a few bad actors".53
- 29 SEPOpenAIOpenAI announces a Decisions API, in limited preview, at its DevDay. Almeida posts "begun, the clone war has".54
- 1 OCTAWS and Cloudflare
OpenAI says the API puts its Luna model to work on user-defined questions with a fixed set of answers. Its API changelog had no entry for it on 3 October.54 Almeida followed his joke with "jk, I love openai", and said he hoped it was a sign that "building in a system one compatible way is the future".54
Strands Decider
Strands Decider 2B is built on Qwen3.5-2B and was released as open source, with its training data and scripts.55 Its authors removed the language-model head, which takes away the model's ability to generate text. A pointer head of just over a million parameters scores each option in its place, and the base model is tuned through a rank-16 LoRA adapter.55
Its median latency is about 115 ms on an RTX 3090, and about 153 ms for small tasks on an M3 MacBook.55 On accuracy and calibration its authors place it "3rd of 33 in the 2B class" on JevBench, a third-party ranking of decision models.55
Clef
Cloudflare's Clef freezes Qwen3.8-27B, and Clef-flash freezes Qwen3.5-9B. Each trains a routing head along with rank-256 adapters.56 The training loss pairs label-smoothed cross-entropy with a Brier loss for calibration. Clef takes images, which Jev does not, and both models are released under an Apache 2.0 license.56 Cloudflare calls them "fully Jev-API compatible". It also says it "developed Reinforcement Learning for Calibrated Decisions (RLCD)" as a secondary training objective, using the name TypeSafe gave its own method.56
What is new
Kadavath's team reported in 2022 that larger language models are well calibrated on multiple-choice and true-or-false questions.3838 Temperature scaling and the reject option for deep networks both date from 2017.39393737
TypeSafe's API takes typed questions, many to a request, and charges nothing for the answers.2424 The model behind it was trained for that format, with a method TypeSafe has not published.2222
The version serving today is jev-1.13.0, and the aliases jev-latest and jev-preview both point to it.2424 Almeida told Latent Space that TypeSafe will not change a model once it is deployed, but is "not promising long-term support" for its versions. He said TypeSafe might keep 1.13.0 for a while, because so many developers use it.2525
Notes
-
TypeSafe, "Introducing System One Models & Jev", launch post by Diogo Almeida, 15 September 2026: "After two years in stealth"; definition; comparison table (LLM input "from $0.20 to $10 / MTok", output "~5x more expensive than input tokens"; Jev output "FREE (too cheap to meter)"; "70ms-500ms" end to end; "40x-200x faster"); RLCD; evals disclosure, including the "System One LLM wrapper", which "tends to be slower and more expensive than giving decisions without probabilities", and reference labels that bias answers "towards OpenAI and Anthropic's models"; hallucination chart note; FAQ on the names, on whether Jev is "just a smaller LLM", on public benchmarks and on training data. https://typesafe.ai/blog/introducing-system-one-models-and-jev 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28
-
DCVC, "TypeSafe emerges from stealth with a new way of doing AI", James Hardiman, 15 September 2026: "DCVC leads the $40 million Series Seed"; latency "less than 100 milliseconds". https://www.dcvc.com/news-insights/typesafe-emerges-from-stealth-with-a-new-way-of-doing-ai/ 2
-
Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026. https://news.ycombinator.com/item?id=49718437
-
Daniel Kahneman, Thinking, Fast and Slow, Penguin Random House listing (published 25 October 2011), the page TypeSafe links. https://www.penguinrandomhouse.com/books/89308/thinking-fast-and-slow-by-daniel-kahneman/
-
TypeSafe press release, Business Wire, 15 September 2026, read in the verbatim copies on Morningstar and Yahoo Finance: "Founded in 2024 and headquartered in San Francisco"; "a nod to Jevons Paradox"; "less than 100 milliseconds of latency"; "up to 100 times faster and less expensive". https://www.morningstar.com/news/business-wire/20260915525333/typesafe-ai-emerges-from-stealth-with-40m-in-funding-with-new-model-for-composable-ai and https://finance.yahoo.com/technology/ai/articles/typesafe-ai-emerges-stealth-40m-190000776.html 2 3 4
-
TypeSafe team page, undated, read on 3 October 2026. https://typesafe.ai/team 2
-
Ouyang et al. (2022), "Training language models to follow instructions with human feedback", arXiv 2203.02155, 4 March 2022. Almeida is the 4th of 20 authors. https://arxiv.org/abs/2203.02155
-
Christiano et al. (2017), "Deep reinforcement learning from human preferences", arXiv 1706.03741, 12 June 2017. Almeida is not among the six authors. https://arxiv.org/abs/1706.03741
-
TypeSafe docs, API reference, read on 3 October 2026: one endpoint; up to 255 options per Choice, each with an optional description; at least two and up to 10 levels per Score; the Choice example on the payouts ticket (0.88, 0.12, 0.0, confidence 0.81, 318 input tokens) and the Score and Noul examples on the same ticket; the four example responses on the page report 18 to 34 output tokens. https://docs.typesafe.ai/api 2 3 4 5 6 7 8 9 10 11 12
-
TypeSafe docs, Primitives, read on 3 October 2026: question IDs "are not sent to the model"; "One question's answer is not hidden context for another". https://docs.typesafe.ai/primitives 2
-
Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026: "noul" is short for bernoulli and maps to if-statements, choice maps to a match statement and score maps to sorting. https://news.ycombinator.com/item?id=49718407 2
-
TypeSafe docs, Confidence, read on 3 October 2026. Choice: (p_max - 1/n) / (1 - 1/n). Score: 1 minus the probability-weighted distance from the most likely level, divided by the same measure for a uniform spread, floored at 0. Noul: the docs suggest |2p - 1|. https://docs.typesafe.ai/confidence 2 3 4 5
-
Arithmetic on the docs examples. Choice: (0.88 - 0.333) / (1 - 0.333) = 0.547 / 0.667 = 0.82, printed as 0.81. Score: 0 x 0.0 + 1 x 0.95 + 2 x 0.05 = 1.05; confidence 1 - 0.05 / 0.667 = 0.925, printed as 0.92. https://docs.typesafe.ai/api and https://docs.typesafe.ai/confidence 2
-
TypeSafe docs, Introduction, read on 3 October 2026: questions evaluated "in parallel and in isolation"; adding questions "barely changes the response time" and "does not create context-rot". https://docs.typesafe.ai/introduction 2 3
-
OpenAI API changelog, entry of 6 August 2024: "Launched Structured Outputs", read on 3 October 2026. https://developers.openai.com/api/docs/changelog
-
Simon Willison, post on OpenAI's Structured Outputs, 6 August 2024. He writes that OpenAI is "presumably" using the same trick as jsonformer and llama.cpp grammars, which interact with next-token selection so that only tokens matching the schema are chosen. https://simonwillison.net/2024/Aug/6/openai-structured-outputs/
-
Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026: constrained decoding makes models "dumber", because "simply masking logits is insufficient": a model that assigns probability to an invalid token "is by definition confused". https://news.ycombinator.com/item?id=49718849
-
TypeSafe home page and FAQ, undated, read on 3 October 2026. https://typesafe.ai/ 2 3 4 5 6
-
Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026. https://news.ycombinator.com/item?id=49719122
-
TypeSafe docs, Parallel questions cookbook, read on 3 October 2026: 13 questions over the GDPR article, asked in one call and as 13 single-question calls, "12.2x cheaper and 10.0x faster with no change in answers"; "The document dominates every request", so 13 calls pay for it 13 times, in 13 round trips. The Primitives page describes the same cookbook as 11.5x cheaper and 9.6x faster. https://docs.typesafe.ai/cookbooks/parallel_questions and https://docs.typesafe.ai/primitives
-
TypeSafe docs, Machine learning primer, read on 3 October 2026: RLCD after RLHF and RLVR; calibration as outcomes given 0.2 happening about 20% of the time. https://docs.typesafe.ai/introduction/machine-learning-primer 2
-
No RLCD paper, method description, model size, base model, revenue or cost figure appears on TypeSafe's site, blog or docs as read on 3 October 2026. https://docs.typesafe.ai/llms-full.txt 2 3
-
Diogo Almeida, comment in the Hacker News launch thread, 15 September 2026: "architecture is close to the chest for now, but we have talked about writing a paper". https://news.ycombinator.com/item?id=49718824
-
TypeSafe docs, Models, read on 3 October 2026: price $42 per Btok and $0.042 per Mtok, "Charged per input token. Output tokens are free."; rate limits; "64k tokens per request; 32k tokens for state plus the longest question"; text-only input; languages; current model jev-1.13.0 with both aliases; not fine-tuned on customer data, the same weights for every account. https://docs.typesafe.ai/models 2 3 4 5 6 7 8 9 10
-
Latent Space, interview with Diogo Almeida, published 21 September 2026, with transcript: synthetic data at 00:22:15; a trillion tokens a day, "even at night", at about 00:35:32; versions, and a possible temporary long-term support for Jev 1.13.0, at about 00:48:11. https://www.latent.space/p/jev 2 3
-
Arithmetic on TypeSafe's prices: $0.20 / $0.042 = 4.76; $10 / $0.042 = 238.1; 238 x $0.042 = $9.996. https://typesafe.ai/blog/introducing-system-one-models-and-jev and https://typesafe.ai/ 2
-
Arithmetic at list prices, input only: 318 tokens x 1,000,000 tickets = 318 million tokens; 318 x $0.042 = $13.36; 318 x $0.20 = $63.60; 318 x $10 = $3,180. https://docs.typesafe.ai/api and https://typesafe.ai/blog/introducing-system-one-models-and-jev
-
Arithmetic on the published limits: 100,000 tokens x 86,400 seconds = 8.64 billion tokens a day; 8.64 x $42 = $362.88; 40 requests x 86,400 = 3,456,000 requests a day; 40 x 318 tokens = 12,720 tokens a second, under the 100,000 token cap. https://docs.typesafe.ai/models 2
-
OpenRouter, Jev Latest listing, read on 3 October 2026 ("32,000 token context window"); Cloudflare, Clef announcement, 1 October 2026 ("our model has a 64k context window (compared to Jev's 32k)"). https://openrouter.ai/~typesafe/jev-latest/ and https://blog.cloudflare.com/clef-decision-models/
-
TypeSafe, Workflow evals, read on 3 October 2026. Each point averages a setup's accuracy, cost and time per item over four workflows with equal weight. The chart prints short labels without vendor or version: Jev 67.8%, $0.0004, 0.4 s; sol 74.1%, $0.0836, 23.3 s; opus 5 73.1%, $0.1761, 37.8 s; terra 67.9%, $0.0304, 10.1 s; sonnet 5 67.8%, $0.1174, 78.1 s; luna 66.8%, $0.0033, 12.9 s; DS v4 pro 65.5%, $0.0413, 86.5 s; DS v4 flash 64.4%, $0.0059, 51.9 s; haiku 4.5 53.6%, $0.0195, 12.5 s. By workflow (Security Incidents, Agent Trace Observability, Invoice Processing, Customer Service), Jev scores 61.7, 71.6, 61.8 and 76.0, sol 62.5, 76.6, 79.1 and 78.3. The chart also plots each LLM run from a standalone prompt; those runs score lower and are left out here. https://evals.typesafe.ai/ 2 3 4 5
-
Arithmetic on the eval values. Pareto frontier, by cost and by time alike: Jev (67.8%), terra (67.9%), sol (74.1%); every other setup is beaten on both counts by one of these. Terra against Jev: 0.0304 / 0.0004 = 76; 10.1 / 0.4 = 25.3. Cost: Opus 5, 0.1761 / 0.0004 = 440. Time: Sonnet 5, 78.1 / 0.4 = 195; Opus 5, 37.8 / 0.4 = 94.5. Means of the eight LLM setups: cost 0.4775 / 8 = 0.0597, and 0.0597 / 0.0004 = 149.2; time 313.1 / 8 = 39.1, and 39.1 / 0.4 = 97.8. Jev's $0.0004 is rounded, so the cost ratios are approximate. https://evals.typesafe.ai/ 2 3 4 5
-
TypeSafe docs, Use-case map, read on 3 October 2026. https://docs.typesafe.ai/concepts/use-case-map
-
Post by Hassan (@nutlope) on X, 17 September 2026, read as quoted in Diogo Almeida's post. Each paper was summarized with DeepSeek V4 Flash before its title, summary and 24 possible topics went to Jev. https://x.com/CompleteSkeptic/status/2100443429281120508
-
Guillermo Rauch, post on X, 16 September 2026: the fx reviewer "runs on GPT Luna today". https://x.com/rauchg/status/2100307962262872105
-
Ibrahim and Zaki, "Evaluating Decision Models for Text Annotation in Computational Social Science", arXiv 2609.24574, 21 September 2026 (v1), preprint. The abstract calls Jev "the first commercial decision model"; the full text names it. Figures from the abstract and the calibration table: Jev's median expected calibration error 0.157, against 0.066 for Opus 5, 0.073 for Fable 5.1 and 0.075 for Sonnet 5, whose confidence is verbalized; the abstract prints the pair as "0.157 against 0.066". Median accuracy 0.815 at or above 0.9 confidence; routing low-confidence items to an LLM "matches or exceeds the LLM alone at a quarter to half of its cost". https://arxiv.org/abs/2609.24574 and https://arxiv.org/html/2609.24574 2 3 4 5
-
TypeSafe docs, System One, read on 3 October 2026. https://docs.typesafe.ai/concepts/system-one
-
Geifman and El-Yaniv (2017), "Selective Classification for Deep Neural Networks", arXiv 1705.08500, 23 May 2017: the classifier "rejects instances as needed, to grant the desired risk". https://arxiv.org/abs/1705.08500 2
-
Kadavath et al. (2022), "Language Models (Mostly) Know What They Know", arXiv 2207.05221, 11 July 2022. https://arxiv.org/abs/2207.05221 2
-
Guo, Pleiss, Sun and Weinberger (2017), "On Calibration of Modern Neural Networks", arXiv 1706.04599, 14 June 2017. Temperature scaling is described as "a single-parameter variant of Platt Scaling". https://arxiv.org/abs/1706.04599 2
-
No expected calibration error, reliability diagram, Brier score or log-loss for Jev in TypeSafe's full docs or launch post, searched on 3 October 2026. https://docs.typesafe.ai/llms-full.txt
-
TypeSafe docs, Classification using confidence cookbook, read on 3 October 2026: 60 filings, a cutoff of 0.9, the confident half right "90% of the time", the other half "40%", so 27 of 30 and 12 of 30; "Numbers below came from jev-1.12 on 2026-08-12". https://docs.typesafe.ai/cookbooks/classification_using_confidence 2
-
Rafe and Das, "Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)", arXiv 2609.24052, 21 September 2026, preprint. Figures from the abstract: F1 0.908; "One frontier model gains 0.059 and the other is indistinguishable from it"; "Recalibration on the same labels reduces calibration error by a factor of 3.3". The full text fits the recalibration map to Jev's probabilities. https://arxiv.org/abs/2609.24052 2
-
TypeSafe docs, Model jaggedness, jev-1.13, "Last reviewed 2026-09-17", read on 3 October 2026. https://docs.typesafe.ai/model-jaggedness/jev-1.13 2 3 4 5 6
-
Sun and Xu, "Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It", arXiv 2609.26758, 22 September 2026 (v1), preprint, figures for the hosted model from the abstract. https://arxiv.org/abs/2609.26758 2
-
Diogo Almeida, comments in the Hacker News launch thread, 15 September 2026. https://news.ycombinator.com/item?id=49718767 and https://news.ycombinator.com/item?id=49719080
-
TypeSafe docs, Coding agents, read on 3 October 2026. https://docs.typesafe.ai/introduction/coding-agents
-
Arithmetic, illustrative only: 10^12 tokens x $42 per 10^9 tokens = $42,000 a day, if every token were billed at list price. The date of the milestone and the share of paid traffic are unknown. https://www.latent.space/p/jev and https://docs.typesafe.ai/models
-
Hacker News, launch thread, item 49717558, posted 15 September 2026; 1,989 points and 520 comments on 3 October 2026. https://news.ycombinator.com/item?id=49717558
-
Diogo Almeida on X, 17 September 2026. https://x.com/CompleteSkeptic/status/2100454726462804333
-
Vercel on X, 18 September 2026: about 13% of teams in the first day, in Vercel's wording, which does not say paid teams. https://x.com/vercel/status/2101077346203971900
-
TypeSafe on X, 20 September 2026. https://x.com/typesafeai/status/2101786156572823624
-
TypeSafe on X, 22 September 2026. https://x.com/typesafeai/status/2102281508950307159
-
Diogo Almeida on X, 27 September 2026 ("ANYONE CAN SIGN UP"). https://x.com/CompleteSkeptic/status/2104338649999626397
-
OpenAI, DevDay 2026 recap, 29 September 2026; the page returned HTTP 403 on 3 October, and its indexed text describes the Decisions API as focusing Luna "on a specific set of user-defined questions with finite pre-defined answers", in limited preview. OpenAI Developer Community, "DevDay 2026 announcements and developer resources", 29 September 2026: "Decisions API is in limited preview". Almeida's post on X, 29 September 2026. OpenAI API changelog, read on 3 October 2026, with no Decisions entry. https://openai.com/index/devday-2026-recap/ and https://community.openai.com/t/devday-2026-announcements-and-developer-resources/1402006 and https://x.com/CompleteSkeptic/status/2105000685209313736 and https://developers.openai.com/api/docs/changelog 2 3
-
Marc Brooker, Mike Chambers and Fabio Nonato de Paula, "Introducing Strands Decider 2B", Strands Agents blog (AWS), 1 October 2026: the LM head removed, "taking away its ability to generate text"; a pointer head of "just over a million total parameters"; a rank-16 LoRA; about 115 ms median on an RTX 3090 and about 153 ms for small tasks on an M3 MacBook; "3rd of 33 in the 2B class" on accuracy and calibration. https://strandsagents.com/blog/introducing-strands-decider/ 2 3 4 5
-
Michelle Chen, Alex Reneau and Kevin Flansburg, "Introducing Clef: our open-source decision models, and new RL fine-tuning platform", Cloudflare blog, 1 October 2026. https://blog.cloudflare.com/clef-decision-models/ 2 3 4