메뉴
HN
Hacker News • 10일 전

제브(Jev): 기존 프런티어 모델보다 40~400배 저렴하고 20~200배 빠른 신규 모델 공개

IMP
8/10
핵심 요약

TypeSafe AI가 2년간의 비공개 개발 끝에 소프트웨어가 직접 사용할 수 있는 빠르고 구조화된 의사결정에 특화된 새로운 모델 클래스 'System One Model'과 첫 공개 모델 'Jev'를 출시했다. Jev는 기존 LLM과 비슷한 수준의 지능을 유지하면서도 두 자릿수 배수(약 100배) 빠르고 효율적이며, 문자열 생성을 포기하는 대신 구조화된 출력에 최적화되어 환각(hallucination)이 원리적으로 불가능하다. 가격 투명성과 검증 가능한 성능 주장을 내세워 업계의 주목을 받고 있다.

번역된 본문

TypeSafe, System One 모델 및 Jev 발표 (자세히 읽기: TypeSafe AI 선언문, 팀 소개, 대기자 등록)

회사 소식, 2026년 9월 14일: System One 모델 및 Jev 소개 작성자: Diogo Almeida, TypeSafe 창립자

모델은 수년 전부터 채팅에서는 이미 초인적인 수준이었습니다. 그렇다면 자동화는 어디에 있을까요? 이것이 저를 지난 4년간 이끌어온 핵심 질문이었습니다. 저는 OpenAI에서 언어 모델이 지시를 따르고 사람과 대화하는 데 유용해지도록 만든 방법론을 구축하는 일을 도왔고, 그 연구가 결국 ChatGPT의 기반이 되었습니다. 당시에는 채팅 모델이 AGI로 이어질 수도 있다고 생각했지만, 아무리 화려한 기대에도 불구하고 정말 중요한 무언가가 빠져 있다는 것이 저에게는 분명해졌습니다.

2년간의 은밀한(스텔스) 개발, 수많은 기술적 난제, 그리고 연구적 돌파구 끝에... 오늘 TypeSafe AI가 첫 번째 System One Model을 출시하게 된 것을 정말 흥분된 마음으로 발표합니다. System One Model은 소프트웨어가 직접 사용할 수 있는 빠르고 구조화된 의사결정을 내도록 설계된 새로운 클래스의 프런티어 모델입니다. 우리는 자동화에 전적으로 집중한 완전히 새로운 스택을 구축했습니다. 새로운 모델 아키텍처, 최대 효율을 위한 병렬 샘플러(parallel sampler), 그리고 '보정된 의사결정을 위한 강화학습(Reinforcement Learning for Calibrated Decisions, RLCD)'이라고 부르는 훈련 방법입니다.

첫 공개 모델은 Jev이며, 오늘 얼리 액세스로 이용할 수 있습니다. Jev는 System One 과제에서 기존 LLM과 비슷한 수준의 지능을 달성하면서도 두 자릿수 배수(약 100배) 더 빠르고 효율적입니다. Jev는 문자열 생성을 포기하는 대신 구조화된 출력에 최적화되어 있으며, 환각(hallucination)을 일으킬 수 없습니다. Jev를 '프런티어 지능 함수 호출'이라고 생각하시면 됩니다. 비구조화된 상태(state)가 입력되고, 타입이 지정된 확률적 의사결정이 출력됩니다.

놀라운 주장에는 놀라운 증거가 필요하니, 아래에서 그 증거들을 확인하세요.

기존의 그리고 새로운 프런티어 증거 / 기술적 결과

우리는 회의론자를 사랑하며, 우리 자신도 회의론자입니다. 쉽게 검증할 수 있는 주장들이 있습니다:

  • 호출당 속도: 우리는 정말로 그만큼 빠릅니다. 다만 공개된 평가는 대부분 미 서해안(현재 서비스가 위치한 곳)의 노트북에서 실행되었습니다.
  • 호출당 비용: 가격은 투명하게 공개합니다. 보조금을 받는 것이 아님을 증명할 수는 없으며, 가격의 지속 가능성은 장기적으로 입증해야 합니다(가격은 오르기보다 내려갈 것으로 예상합니다).
  • 타입 오류 없음: 단 하나의 반례만으로도 쉽게 반증할 수 있지만, 수학적으로 불가능합니다.

더 대담한 주장에 대해서는 가능한 한 많은 세부 맥락을 제공하고자 합니다.

나란히 비교 시연 우리의 나란히 비교 데모는 우리 모델과 LLM의 핵심적 차이를 보여줍니다. Jev는 토큰별로 자기회귀적으로 생성하는 대신, 모든 확률을 병렬로 출력합니다. 문자열은 매우 강력하고 범용적이지만 비용이 큽니다. 문자열을 '포기'함으로써 오히려 많은 초능력을 얻었습니다!

세부 사항: TypeSafe 얼리 액세스 사용자를 위해 실제 쿼리를 공개합니다. 이 쿼리는 매우 단순화되었으며, 화면의 출력을 이해하기 쉽도록 설명적이고 사람이 읽을 수 있는 키를 가진 질문들을 선택했습니다. 상태(state)도 샘플링 방법론의 차이를 강조하기 위해 짧고 밀도 높고 상세한 문단으로 구성했습니다. 비교적 짧은 입력은 우리 모델에 유리하게 작용합니다. 눈썰미가 좋은 분이라면 알아차렸겠지만, 녹화된 실행에서 GPT-5.6 Terra와의 유일한 불일치는 '이탈 가능성 수준(Churn likelihood level)'입니다. 실제 정답은 우리에게도 정말 모호해 보입니다. 이 예시에서는 기본 추론 설정의 GPT-5.6 Terra를 사용했는데, 평균적으로 Jev와 지능 면에서 가장 비교 가능한 모델로 판단했기 때문입니다. 재미있는 사실: 비슷한 데모가 바로 우리가 System One Model 방향에 전력투구하기로 결심한 계기였습니다!

워크플로 평가 우리는 AI가 코드 안에서 얼마나 잘 작동하는지 측정하는 새로운 유형의 평가를 만들었습니다. 우리는 정답 분류에 최적화하지 않으며, 테스트 하네스(harness)와 모델이 변경되는 것을 허용합니다(하네스 엔지니어링을 통한 과적합 가능성도 열어둡니다). 대신 올바른 연산 그래프(코드로 표현된 '워크플로')가 존재한다고 가정하고, 가장 크고 똑똑하고 비싼 모델의 예측을 사용합니다...

원문 보기
원문 보기 (영어)
TypeSafe announces System One models and Jev Read More TypeSafe AI Manifesto Our Team Join Waitlist TypeSafe announces System One models and Jev Read More TypeSafe AI Manifesto Our Team ∵ Back Company News Sep 14, 2026 Introducing System One Models and Jev Diogo Almeida, founder, TypeSafe Models have been superhuman at chat for years, so where is all the automation? This has been my driving question for the last four years. At OpenAI, I helped build the methods that made language models useful at following instructions and talking with people. That work ended up as the research behind ChatGPT. At the time, I thought maybe chat models would lead to AGI, but despite the hype it became obvious to me that there was something really big missing. After two years in stealth, countless technical challenges, and research breakthroughs… I am beyond excited to announce that today, TypeSafe AI is releasing our first System One Model : a new class of frontier models built to make fast, structured decisions that software can use directly. We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD). Our first public model is Jev , available today in early access. Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient. While Jev gives up string generation, it’s optimized for structured outputs and can’t hallucinate. Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out. Extraordinary claims require extraordinary evidence so see below for the receipts. 💅 Frontiers, Old and New Evidence / Technical Results We love skeptics, and are skeptics ourselves. There are some claims you can easily verify: Speed per call: We truly are that fast, though our published evals are generally run from our laptops on the West Coast (this is where our service is currently based). Cost per call: We make our pricing transparent. We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up). No type errors : This would be an easy thing to falsify with just a single counter-example, but it is mathematically impossible. For our bolder claims, we want to provide as much nuance as we can. Side-by-side demonstration Our side-by-side demo shows a key difference between our models and LLMs: Jev outputs all probabilities in parallel instead of autoregressively generating by token. Strings are extremely powerful and general, but costly. “Giving up” strings actually gives us a lot of superpowers! Nuance: For people with early access to TypeSafe, here is the actual query . The query is highly simplified and questions were chosen to have descriptive, human-readable keys so that the output on the screen is understandable. The state is also a short, dense, and detailed paragraph, to emphasize the difference in sampling methodology. The relatively shorter input paints our model in an advantageous light. For the keen eyed, for the recorded run, the only disagreement with GPT-5.6 Terra is on “Churn likelihood level”. The actual answer seems genuinely ambiguous to us. We used GPT-5.6 Terra with default reasoning for this example, because we’ve found it to be the most comparable at intelligence to Jev on average. Fun fact: a similar demo was what convinced us to go all-in in the direction of System One Models! Workflow evals We made a new type of evaluation to measure how well AI works within code. We don’t optimize for a ground truth classification orand allow the harness and model to change (potentially allowing for overfitting via harness engineering). Instead, we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable). Jev is off the charts – owning the Pareto frontier for almost 2 orders of magnitude. We also compare to models with a generated prompt doing all the logic in their chain-of-thought, but this tends to do significantly worse than using the workflow itself. Note that the calls here are significantly more complex than the side-by-side demonstration above. That’s because they’re more representative of the types of production workloads needed for true business automation. Below is the simplest of the 4 workflows we’re publishing: The most reliable real-world workflows tend to have many independent, decomposed questions, with fine-grained behavior that’s dependent on probabilities instead of discrete decisions. The end result is discrete branching, but how we get to a final answer involves a lot of domain-specific engineering that needs to be done highly consistently. See our workflow evals site for all the details: examples, disagreements, full queries, and each workflow. Nuance: This is where the claims of 193.6x faster, 444.6x cheaper on our home page comes from, and we expect that these are on the higher end of real world gains. These content of these workflows were not deliberately chosen nor constructed to make our model look good, and are not in our training distribution. However, they were made by individuals on our model capabilities team, so some bias could exist. We use the average of GPT-6 Astra and Fable 5.1 as the reference answer, which biases answers towards OpenAI and Anthropic’s models. We likely underestimate the relative performance of our model and DeepSeek’s models. The LLMs use our System One LLM wrapper, which constrains LLMs to output structured decisions compatible with our API. We have found this to be the most accurate way to get decisions from LLMs, but this tends to be slower and more expensive than giving decisions without probabilities. Hallucination and Type-safety Hallucination and type-safety are intrinsically related, and we think the latter is table stakes for automation. Having a hallucinated tool call is inconvenient in an agent, but is an absolute deal-breaker if it’s part of a system with latency guarantees or it’s buried several layers deep in a dependency chain. Existing models, no matter how smart , still hallucinate and have type errors. Nuance: The numbers for LLMs are from OpenRouter i.e., there almost certainly is bias here: more complex queries might be routed to better models. Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots. Fun Demos Perhaps the most exciting part of our work is enabling new use cases. We have a lot more to show you, but here are a couple of the team’s favorites: Doom We love how this doomo doomonstrates real-time intelligence and what can be doone with code + AI. The engineer behind it was worried about making 10 queries a second (which ends up costing ~$7/hour), but the rest of us agreed that was lower than expected! This is so fun we intend to not only release an in-depth walkthrough, but also host some events to hack on this. Nuance: The demo is on structured state as a data structure with text, not on images (yet…) A non-AI doom bot could play better, but we wanted a bot that was reactive to different representations of game state, and most importantly… following instructions was cool as heck! Wikiracing The objective of the game is to start on one Wikipedia page and reach a specific other Wikipedia page using only links you come across while traversing. Each step can mean choosing between hundreds to thousands of links! It’s a great playground for demonstrating not just intelligence-per-second, but also the compounding benefits of not hallucinating with high-cardinality choices. Nuance: As far as we know, it was complete