メインコンテンツへ移動 / Skip to main content

Making Jev’s Speed and Cost PlayableTwo Hours of Conversation, Zero Handwritten Lines

Trying Jev turned into a buzzword game through conversations with Claude Code. A real demo, roughly 400ms responses, about ¥5 in API usage, and the mistakes that moved decisions back into code.

The actual game with IT terms filling category meters, cost on the left and API logs on the right
Technology
Published on: October 1, 2026
Read time: 9 min
Author: Pochang Lab
Read time: 9 min

0. Background: what Jev is

  • TypeSafe is the name of the service (the provider), and Jev is the name of its model in what it calls the “System One” family. In the API you call it by specifying something like “jev-latest”. [4]
  • The biggest difference from an ordinary LLM is that it does not generate prose. What comes back is only a judgment in a type you define, plus probabilities. [2]
  • There are three question types.
    • Noul: a yes/no judgment. Instead of true/false, it returns the probability (0–1) that the answer is yes. Example: “Is this support request urgent?” → 0.95
    • Choice: picks one option from a set you define, and also returns the probability of every option and a confidence value. A single question can have up to 255 options. Example: “Which field does this term belong to?”
    • Score: returns where something sits on levels you describe in words (up to 10 levels), as a decimal such as 1.43. Example: “How current is this technology?”
  • A request sends the material to judge (state) and the list of questions (questions) once. Questions that are independent of each other can be sent together and are answered in parallel. [3]
  • Pricing applies only to input tokens; output is free. [4]

An LLM can make the same judgments. But it tends to involve listing the options in a prompt, having the model answer in prose, parsing that prose, and—if you want to know how confident it is—asking the same question several times to see whether the answers agree. The extra generated text and round trips add time and cost. It helps to think of Jev as a model narrowed down to the cases where you only want the judgment.

That is the background. From here on, this is the story of actually using it.

1. I wanted to make “apparently cheap” something I could play

Jev came up in conversations at work. My reaction was, “Wait, am I already falling behind?” I looked into it and kept hearing that it was fast and cheap. But calling an API once and staring at a number did not tell me what would be fun to build with it.

What eventually emerged was IT Expert Challenge, a small game where you enter IT terms to fill five category meters. There is no Send button. A judgment comes back as you type, and the word flies into its category. Cost appears on the left; HTTP requests and responses appear on the right. I wanted the speed and cost to be visible as part of a working application. [1]

Watch the sequence from input to judgment to movement into a meter. Demo mode (?demo) types automatically, one character at a time, imitating keystrokes. API response times are measured; cost is calculated from measured token usage. The footage is not sped up; only the first second of the recording was trimmed. The game UI is Japanese.

I had also heard a rumor that signing up came with $5 in free credit. That did not happen with my account. Creating a key required a paid balance, and the minimum top-up was $5, so that was what I added. This describes my experience, not a promise about free credit or minimum purchases for every future user.

2. I had not even planned to build an app

My initial request to Claude Code was simply, “First, make it usable.” During our conversation, the agent produced a small tool to expose what happened inside a call, along with a page for checking it. From there, I mostly used voice input to say things like “more like this” or “what about that?”

The first version was a conversation meter. But talking at it and watching a meter rise was not very entertaining. When we turned it into a game about collecting IT terms, follow-up questions started getting in the way. Rather than asking me to elaborate, it needed to make me want to supply another term. If a recording icon lingered on screen, I would say it looked bad and ask for it to be removed. The project grew through decisions like those.

About two hours of conversation. Zero lines of code written by a human. That does not mean I did nothing. I decided what I wanted to test, what felt fun, and what felt wrong, then changed direction after seeing the result.

All five meters and the cost counter start at zero, above an input field with no Send button

Two terms fill a category; filling all five clears the game. Category names themselves, such as “database,” do not count. Empty meters make it clear what to try next. Click a screenshot to enlarge it.

For this small project, the more implementation the agent handled, the more my work shifted toward choosing the direction and deciding what to try. I did not need a complete design at the start. If I could explain what bothered me about a working version, we could test another shape. That back-and-forth was part of the fun.

3. Give Jev the judgments and code the calculations

The game uses all three question types introduced in section 0: Noul, Choice and Score. [2]

I considered having an LLM write the responses. But then the wait for generated prose would influence how fast the game felt. Since I wanted to feel Jev’s speed, the running game uses only ordinary code and Jev. Claude Code, used during development, had a different role.

  • Jev: judges category, how current a technology is, learning difficulty, how niche it is, and which technology name matches a trivia entry.
  • Code: splits input, extracts candidate terms, and calculates points. It selects reactions from the results and checks for previously collected terms. A local proxy injects the API key; the browser never receives it.

Trivia is also selection, not generation. I prepare lines for Oracle, for example, and ask Jev which technology name the input represents. A distorted spelling such as “morakuru” sometimes still matched Oracle. Jev handles the fuzzy part that an exact-match dictionary would struggle with, while code controls the words that appear on screen.

Oracle is classified as a database term and a prepared remark about licensing cost appears

Jev did not write this sentence. Code used the technology-name match to select one of the prepared Oracle lines.

4. Five questions per word, one request

What category does Oracle belong to? How current is it? Is it difficult? Is it niche? Which trivia entry fits? Instead of making five consecutive calls, I send five questions against the same state in one request. TypeSafe’s documentation describes evaluating independent questions about shared state in parallel. [3]

The JSON structure looks like this. Question wording and options are shortened for explanation; this excerpt is not a complete configuration to run unchanged. The application’s actual prompts are in Japanese.

json
{
  "model": "jev-latest",
  "state": { "segment": "Oracle" },
  "questions": {
    "t0": { "type": "choice", "instructions": "Which category is Oracle in?",
      "criteria": { "database": "Database", "not_it": "Not an IT term", "...": "Other categories" } },
    "n0": { "type": "score", "instructions": "How niche is it?", "criteria": ["Famous", "Common", "Niche"] },
    "e0": { "type": "score", "instructions": "How current is it?", "criteria": ["Past", "Older", "Established", "New"] },
    "h0": { "type": "noul", "instructions": "Is the learning cost high?" },
    "k0": { "type": "choice", "instructions": "Which technology does the name itself identify?",
      "criteria": { "oracle": "Oracle Database", "none": "No match", "...": "Other technology names" } }
  }
}

The Oracle response recorded in the README took 386ms and used 1,642 input tokens. Here is part of its answers object. Rather than an explanation to parse, it contains values code can use directly.

json
{
  "t0": { "type": "choice", "choice": "database", "confidence": 0.99,
    "probabilities": { "database": 0.99, "not_it": 0.01 } },
  "e0": { "type": "score", "score": 1.52 },
  "h0": { "type": "noul", "noul": 0.55 },
  "k0": { "type": "choice", "choice": "oracle", "confidence": 0.92 }
}

Other probabilities and answers are omitted. Code uses the result to add a database point and show an Oracle line. Even with five questions, a response in roughly 400ms makes the gap between input and reaction feel short. The judgments about relevance and difficulty are used to choose game reactions; they are not measurements of technology adoption or learning time.

The right pane shows a five-question HTTP call and typed response, while the left shows average latency and cost

The right-hand log exposes question counts and returned values; the left tracks cumulative cost. I wanted the calls behind the animation to be visible alongside it.

5. About ¥0.15 per game, about ¥5 including the experiments

The clear screen in the featured video reads 25.2 seconds, 14 judgments, 407ms average, and ¥0.1532. The README gives approximately ¥0.14 for another demo record; the footage here shows approximately ¥0.15. The 14 judgments include inputs handled entirely by code. There were 12 actual API calls.

Average latency measures the interval from the proxy sending a request to TypeSafe until it finishes reading the response. It does not include all the typing time and animation. Cost is calculated from the input-token counts in the responses, using a fixed conversion assumption of $1 = ¥160.

The clear screen reports 25.2 seconds, 14 judgments, 407ms average and a cost of 0.1532 yen

The result of the recorded run. The roughly 25-second duration includes automated typing and animation, so it should not be read as API processing time.

The official Models page lists Jev 1.13 at $0.042 per million input tokens, with output free. More input means more cost, but for this game it was low enough that I could keep refining questions without worrying about each individual call. [4]

Including experimentation and repeated demo recordings, the dashboard showed Spend $0.032, Requests 436, and Tokens 840,239, or about 840,000. Converting Spend at the same fixed rate gives ¥5.12, roughly ¥5. This is Jev API usage, not the total development cost including Claude Code or human time.

Usage totals show Spend of 0.032 dollars, 840239 tokens, 436 requests and a Balance of 5.00 dollars

The totals and Balance regions were cropped from the dashboard and stacked vertically. Values are unchanged, and the “Stats may be delayed” notice is retained.

Balance still displayed $5.00. The dashboard explicitly warns that statistics may be delayed, so an unchanged balance does not establish that usage was free. The displayed Tokens figure is also a dashboard total; I did not assume it meant input tokens alone and use it to reverse-calculate Spend.

6. The failures taught me where to draw the line

The most useful lessons came from changing what I asked the model to do.

Repeat detection went back into code. Even when I repeated the exact same sentence, the answer to whether it was a repeat was 0.14. The documented weaknesses include indirection and judgments requiring multiple steps. That single result does not prove its cause, but exact text matching and checking collected terms can be implemented deterministically. There was no reason to delegate those checks to a model. [5]

I spelled out the “not IT” boundary. When asked to choose a category, the model treated the everyday Japanese word for “management” as an IT term with 94% probability. Adding examples such as management, settings, and model to the “not an IT term” description brought that down to 40%. A word appearing in an IT conversation is not necessarily a specialist term. The question needed that boundary made concrete.

Relatedness and identity needed different wording. DynamoDB matched the AWS trivia entry. They are related, but I did not want an AWS line there. I changed the question to ask which technology the name itself identified and added a condition excluding individual services of a broader platform.

More questions did not noticeably increase latency in these trials. Moving from one question to five produced responses mostly around 350–490ms. There were occasional delays of several seconds, however. This is neither a guarantee of 400ms every time nor a claim that question counts can grow without limit.

These observations apply to these inputs and prompts. They do not establish Japanese-language accuracy overall, or show that the same thresholds will work in another application. The implementation specifies jev-latest, so when repeating the experiment I would also record the versioned model ID returned in each response.

7. What a single API call would not have taught me

Just calling an API did not seem very interesting, so I made something I could experience. Low cost became room to revise things repeatedly. Speed became a reaction that encouraged the next word. The failures also exposed which operations belonged in ordinary code.

The interesting part was not only Jev. It was the process of experimenting through conversation with an agent. Even when I was not writing the code, there was still a job to do: decide the direction. What did I want to find out, and what felt wrong? Communicating those things moved this little game forward.

The code and setup instructions are in Jev Buzzword Rush on GitHub. It is a personal experiment. Input is sent to TypeSafe and logged locally, so use text you can safely disclose when trying it. Demo mode also makes real API calls, billed to the owner of the key.

References

  1. [1]Jev Buzzword Rush — README and implementation. The development story is the author’s experience; measurements come from the demo footage, dashboard, and README records. ↩
  2. [2]TypeSafe — Primitives (Questions). Noul, Choice, and Score question and answer types. ↩
  3. [3]TypeSafe — Speculative fan-out. Parallel evaluation of multiple questions against shared state. ↩
  4. [4]TypeSafe — Models. Jev 1.13 input pricing, free output, and model aliases. ↩
  5. [5]TypeSafe — Jev 1.13 jaggedness. Documented limitations involving indirection, comparison, calculation, and precise instructions. ↩

Related Articles

May 16, 2026

Your Home PC Is Becoming a Remote AI Agent Workstation

Using Claude Code Remote Control and Codex mobile access as reference points, this article explains how local development machines are becoming remotely supervised AI agent workstations.

TechnologyRead more
May 12, 2026

What Changed from Claude Opus 4.6 to Claude Opus 4.7?

A practical comparison of Claude Opus 4.7 and 4.6 across coding, agentic work, vision, instruction following, cost, and the broader GPT-5.5 context.

AIRead more
September 30, 2026

Why Orca Passed 80,000 GitHub Stars: One Place to Run, Review, and Return to Parallel Agent Work

A balanced look at Orca’s 81.6k stars: CLI portability, worktrees, review, Design Mode, screen limits, competitors, founders, and the business behind a free agent development environment.

TechnologyRead more
September 4, 2026

Building nagara: A Voice Player for AI Replies, Articles, and Novels

I wanted long Claude Code replies to be easier to listen to. This engineering story follows nagara, an open-source macOS player built with Swift and AivisSpeech, through quiet collection, sentence navigation, and pitch-preserving playback speed.

TechnologyRead more
September 2, 2026

You Don’t Have to Abandon Claude Desktop: The Practical Case for Claude Code CLI, Subagents, and Agent Teams

Subagents work in Claude Desktop's Code tab, so what still makes the CLI different? This practical guide explains CLI-only Agent Teams, Unix pipes, shell-environment continuity, which model and effort level each teammate and Subagent inherits, local-only configuration, Git exclusions, and copy-ready reviewer and log-watcher agents for ordinary Claude subscribers.

TechnologyRead more