awesome-jev
A curated awesome list of public projects and practices built on Jev, TypeSafe AI's System One model for typed decisions.
This README is the homepage aggregate of the current category files, so the latest accepted entries are visible here without drilling into subpages.
Jev is not a chat model. It takes unstructured state plus a typed question and returns a typed decision — a choice, a score, or a boolean, each with a confidence. That makes it a drop-in decision layer for software: classification, routing, rubric scoring, verification, and agent guardrails. This list tracks who is actually building with it, and which patterns transfer across industries.
The repository treats all categories equally — each entry lives in exactly one category, chosen by its direct Jev application domain. A dedicated Related Practices / Discussions category captures credible public practice signals — X threads, Reddit discussions, and interviews — that describe real Jev usage even when no strong standalone case page exists yet.
[!WARNING] A listing is not an endorsement. This project applies inclusion rules only — public, citable, genuinely uses Jev for a typed decision, one-sentence summary. It does not review code quality, security, maturity, or whether a project runs at all.
Treat same-day bulk submissions with particular care. Several repositories published together by one author, sharing a scaffold and a thin commit history, can satisfy every inclusion rule and still be unproven. Volume is not evidence of quality. See Curation is not endorsement for a checklist to run before adopting anything here.
Why this list
Most Jev discussion is scattered across launch threads, model-gateway listings, and one-off prototypes. This list answers two practical questions quickly:
- Where is Jev already making real decisions in production workflows?
- Which decision patterns transfer across industries?
This is not a comprehensive database. It is a high-signal, fast-scanning field guide.
Inclusion criteria
An entry should meet all of the following:
- The source is public and citable.
- The example uses Jev (or a documented Jev port/derivative) for a concrete decision task — not a generic classifier, router, or LLM judge with no Jev involvement.
- The source explicitly names
Jev/jev, cites TypeSafe AI's System One models, or shows a typed-decision loop (typed question → typed answer with confidence → accept/reject/escalate). - The summary explains the scenario, method, and value in one sentence.
We do not include:
- Generic classifiers, routers, or research agents that merely resemble the pattern without using Jev.
- Pure theory or opinion without a concrete practice.
- Launch-hype commentary with no working artifact or reproducible result.
- Long write-ups inside the list itself.
- Sources that are private, inaccessible, or too vague to classify.
Curation is not endorsement
Inclusion means one thing: the entry satisfies the inclusion rules above. It is not a quality review, a security audit, or a recommendation. We do not verify that a project compiles, that its tests pass, that its published numbers reproduce, or that its license permits your use.
This matters most for projects that arrive in bulk. When one author releases several repositories on the same day, they commonly share a single scaffold — the same AGENTS.md, CLAUDE.md, STATE.md, and CHANGELOG.md — land in one or two commits each, and may ship considerably more prose than code. Such projects can be entirely legitimate; they are simply unproven. Treat them as leads, not as validated tools.
Before adopting an entry, check it yourself:
| Check | Why it matters |
|---|---|
| Does the code actually call the Jev API? | An entry can read well on a README alone. Look for a real request carrying typed questions, and a parsed answer coming back. |
| Is there a runnable check? | A test, an example with expected output, or a public demo. No check means no evidence that it works. |
| Do the numbers have a source? | Any accuracy, latency, cost, or volume figure should be traceable to the linked page. We strip claims we cannot verify, but the project page itself may still carry them. |
| How much of the repository is code? | Some projects are mostly prompt documents. That can be legitimate — just know which one you are getting. |
| Is there a license? | A few entries have none, which limits reuse and redistribution. |
Found something wrong? Open an issue or a pull request — removal is as valid a contribution as addition. Rules for AI-assisted work, project depth, and submission rate live in CONTRIBUTING.md.
Current coverage
- Classification & Routing — 62 entries
- Adaptive & Realtime UI — 10 entries
- Verification & Guardrails — 47 entries
- Scoring & Ranking — 40 entries
- Agent Decisions — 60 entries
- Data Labeling & Curation — 10 entries
- Evaluation & Benchmarking — 35 entries
- Calibration & Research — 50 entries
- Infra / SDKs / Integrations — 102 entries
- Game & Simulation — 25 entries
- Robotics & Physical — 9 entries
- Finance & Trading — 8 entries
- Compliance & Legal — 2 entries
- Content Moderation — 8 entries
- Related Practices / Discussions — 102 entries
Open categories still being seeded
- Scientific Pipelines — 0 entries
Each entry lives in exactly one category. When a project could fit multiple categories, we choose the one closest to its direct application domain.
Browse by category
- Classification & Routing (source)
- Adaptive & Realtime UI (source)
- Verification & Guardrails (source)
- Scoring & Ranking (source)
- Agent Decisions (source)
- Data Labeling & Curation (source)
- Evaluation & Benchmarking (source)
- Calibration & Research (source)
- Infra / SDKs / Integrations (source)
- Game & Simulation (source)
- Robotics & Physical (source)
- Finance & Trading (source)
- Compliance & Legal (source)
- Content Moderation (source)
- Related Practices / Discussions (source)
Find by coding agent
Optional tags on an entry name the coding agent it targets and the kind of integration it is. Most entries carry none — they are added only when the source itself supports the classification.
- Multi (22) — jev-router · Harness Router · Switchboard · JevRouter · decision-router · jev-guard (leepokai) · jev-axi · jev-harness · Canny · jev-lint · Perch · Jev-Code-Reviewer · yoshi · public-browser · jev-agent-skill · hermes-jev-skills · Jevbridge · jev-use · jev-mcp (burnigtm) · typesafe-mcp · jev-style · Building with TypeSafe Jev
- Claude Code (15) — jev-skill-router · spending-effort-with-jev · sortwell · jev-seo · Sniff Test · claude-code-templates · jev-secret-guard · jev-retrieval · fast-jev-compaction · jev-belay · jev-pruner · SkillRanker · oh-my-claudecode · jev-opus · jev-auto-approve
- Pi (14) — pi-jev-router · pi-jev-skill-picker · pi-jev · pi-heed · pi-verdict · Reflex · pi-subagent-jev · Jev Deep Research · pi-typesafe-jev · pi-jev (TheoOliveira) · pi-quiet-ask · pi-fast-jev-compaction · pi-typesafe-router · Testing Jev for Pi extensions (r/PiCodingAgent)
- Codex (6) — Codex Jev Router · Jev Auto Router · Foreman · fast-dev-compaction · jev-desktop · jev-browser-use
- DeepSeek Harness (3) — dsh-jev-interceptor · dsh-auto-mode · dsh-jev-decide
- OpenClaw (2) — openclaw-jev-leakguard · openclaw-jev-trigger
- Cline (1) — Cline plugins
Full list
Classification & Routing
Source file: categories/classification-routing.md
- JEV Book Tags
- Library cataloguing: a calibre plugin asks Jev
Noulquestions about book genres and subjects, applies configurable per-tag probability thresholds, and preserves existing tags while leaving uncertain results for review. - Diffusion Jev
- Visual classification: independent Jev-style DiffusionGemma/SGLang server that selects doodle and flower labels from image pixels with typed Choice questions and displays candidate scores in a drawing playground, with public evaluation artifacts and uncalibrated probabilities.
- Notra
- Marketing analytics: production GEO platform whose
NOTRA_JEV_CLASSIFIERSflag routes brand-visibility classifiers off an LLM and onto JevBooleandecisions at a 0.5 threshold, targeting 300 ms p50. - jev-router
- Developer tooling: routes Claude Code tasks to the cheapest capable model by asking Jev to choose among candidates.
- jev-router (prismhq)
- LLM infrastructure: open-source LiteLLM-based router where a Jev decision picks which model serves each request.
- pi-jev-router
- Coding agents: adds automatic per-request model routing to the Pi coding agent through Jev decisions on Vercel AI Gateway.
- jcm-router
- Coding agents: local proxy that picks the Claude model and reasoning effort per message with a Jev decision while leaving the cached main chat untouched.
- Codex Jev Router
- Coding agents: asks Jev Choice and Noul questions about a short task summary to select a Codex subagent model and reasoning effort, with confidence thresholds and a Sol fallback.
- Jev Auto Router
- Coding agents: per-call Codex GPT routing where Jev makes one typed Choice over host-available (model, effort) pairs; a local Responses proxy keeps the tool loop continuous, then independent verification and Router Compass record whether the task still passed (prototype).
- Harness Router
- Coding agents: framework-agnostic tool router that keeps obvious calls on a fast path, uses Jev for genuinely ambiguous choices, and can apply bounded MCTS when multi-step consequences matter, with MCP plus Codex skill and hook integrations.
- GoEventBus
- Event infrastructure: high-performance Go event bus whose optional routing layer evaluates deterministic rules, reuses cached decisions, then asks Jev a typed
Choiceto select the event projection only for ambiguous cache misses. - jev-agent-skill-router
- Agent infrastructure: routes agent skill selection through typed, confidence-aware Jev decisions so weak matches are declined instead of guessed.
- typesafe-jev CV screener
- Recruiting: screens a folder of CVs with Jev typed judgments against an editable policy, re-scoring candidates for free when the policy changes.
- Jev email intent workflow
- Back-office automation: async LangGraph workflow gets a typed Jev
Choice(invoiceorgeneral) and routes each inbound email to the matching handler. - DiffJury
- Code review: routes each pull request by risk with Jev before a human reviewer is assigned, doubling as a review coach.
- HA-Jev
- Smart home: Home Assistant integration that answers questions about the house as a probability, a choice, or a score.
- secondlayer
- Fault triage: self-hosted Stacks data service whose Slack gate and fault-triage paths both run on Jev decisions.
- jev-logtriage
- On-call operations: batches collapsed Loki logs into one Jev call of Noul, Score, and Choice questions, then maps answers in code to suppress, watch, review, notify, or page, with low confidence going to review and nothing executed.
- new-api-typesafe-plugin
- LLM gateway: adds a native
/v1/systemoneendpoint to new-api so typed decisions sit behind the same gateway as chat models. - duet-agent
- Agent harness: keeps a Jev-backed routing table for deciding which model should serve a request.
- omo-jevlike-router
- Skill routing: shrinks the skill catalog in a system prompt with one forward pass over a frozen Qwen, routing each request Jev-style.
- jev-cookbook
- Developer education: 15 runnable Node recipes that route support tickets, file documents, categorize bank transactions and label Gmail with Jev
ChoiceandNoulquestions, sending low-confidence answers to human review. - flue-jev-demo
- Agent routing: routes a Flue agent's work with Jev through Cloudflare AI Gateway.
- DocJev
- Document pipelines: LlamaIndex's open-source library that classifies a document against natural-language category rules or finds the boundaries between sub-documents, with swappable OCR backends (liteparse or LlamaParse) and a benchmark harness whose 40-document pilot classified 40/40 originals correctly at about 182 ms Jev decision p50.
- jev-fit - Developer tooling: hosted fit checker that sends a pasted software idea and a fixed typed rubric to Jev in one call, where a
Choicepicks plain code, Jev or a reasoning LLM behind aNoulgate for non-tasks, code vetoes Jev when the idea needs images, and low confidence returns "not sure"; closed source, free page and API. - jev-skill-router
- Coding agents: Claude Code plugin whose UserPromptSubmit hook asks Jev one
Choiceover the installed skill roster plusBoolean-style gates on whether any skill is needed, suggests a skill only when the gate and the per-candidate fit both clear 0.30, and defaults to a shadow mode that logs the decision without injecting it. - Jev Wrapped
- Media analysis: reads up to 1,500 posts from the last year of a public Telegram channel and asks Jev a
Choiceover ten kinds of post plus threeNoulquestions (paid ad, clickbait, emotional pressure) about each, counting an ad from 0.7, or from 0.4 when the kind is also ad, and clickbait and pressure from 0.5, then draws the monthly mix on a shareable card that links the highest-scoring posts for a manual check. - Jev-Mail
- Email productivity: runs a 24/7 Gmail classifier on user-owned Google Apps Script where Jev scores urgency, importance, and category, routing uncertain or suspicious mail to Review without a local daemon.
- AI-decision-maker
- Data cleaning: asks Jev
Choicequestions to classify CSV columns into a 13-code type vocabulary and each dataset into one of six scenes, then executes every write locally; measured Jev at 6.6–12.7× an LLM's token cost on this task because the output is already one character while per-question criteria repeat. - hearth-jev-rental-search
- Housing search: autonomous multi-source rental search where Jev decides which listings match the criteria.
- pi-jev-skill-picker
- Coding agents: ranks the Pi agent's installed skills against the current task with Jev before any of them run.
- Jevonian
- Coding agents: local OpenAI/Anthropic-compatible proxy where one Jev call picks both the model route and the thinking level for
jevonian/autofrom session state, quota health, candidate capabilities, and cache-switch penalties, after deterministic code has filtered candidates and while pinned models, explicitjevonian/requests, androuting.mode: "off"skip Jev entirely;minConfidencemarks a low-confidence route in the ledger rather than accepting it, and the ledger records the serving model and why. - Switchboard
- Coding agents: open-source System One-powered router that automatically matches each Claude Code or Codex task to an appropriate model and reasoning effort, then keeps that choice stable for the conversation to preserve prompt-cache continuity; powered by Jev today, with Laya, Kev, and Cua-S1 coming soon.
- Tab Sorter
- Browser tooling: Chrome MV3 extension that groups every tab in the window into named, colored Chrome tab groups from one parallel Jev call — one
Choicequestion per tab against user-editable group criteria — with unclassified tabs falling to a fixed fallback bucket, manual groups untouched, and the pre-grouping tab order restored on ungroup; an optional LLM engine with self-invented group names is included for comparison runs. - Feed Lens
- Social media: uses Jev
Nouljudgments against per-platform, user-defined topic and expression labels to annotate Weibo, Threads and X posts directly in a Chrome extension. - jev-table-import-mapper
- Data import: maps an uploaded CSV's columns onto a destination table with a strict deterministic name-equality pass, then one Jev
Noulper remaining (source, destination) pair plus a guardNoulper incoming column, mapping 10 of 10 columns of a 23-column export at 253 questions in one call, 915 ms, $0.0012, unmapped fields left visible above a 0.75 threshold rather than guessed. - jev-oncall
- On-call operations: asks Jev one call of four typed questions per alert (
Noulactionable,Scoreseverity,Choiceowning team,Choiceduplicate-of), pages on P(SEV1)+P(SEV2) ≥ 0.80 and drops only below 0.20 when actionability agrees, sends the band between to a human who has 15 minutes to ack before it pages anyway, links duplicates into a cycle-broken incident graph so two alerts can never silence each other, and falls back to configured severity on any error or timeout; a published 300-alert run measured p50 418 ms, p95 1477 ms and $0.04 per 1,000 alerts, with the slowest call landing 151 ms short of the 2 s timeout. - Jevidence
- Developer education: Python sandbox asks Jev
ChoiceandNoulquestions about issue category and reproduction steps in opt-in live mode, then applies confidence and reproduction gates to propose a queue or review fallback without assigning the issue, with synthetic offline fixtures and policy tests. - JevBystander
- Messaging: Android accessibility app that reads the visible WeChat chat window and answers one batched request of typed
Choice,ScoreandBooleanquestions to sort the peer's message into intent (10 options), an emotion distribution (9 options), urgency (0-3Score) and a suggested reply posture (11 options), then shows exactly three toasts and takes no other action - no generated reply, no input injection, no screenshot or OCR, with a local alias-to-relation table passed as state so the same sentence is judged differently for a partner than for a colleague. - langchain-skill-router
- Agent infrastructure: per-turn skill routing for LangChain deepagents, where Jev ranks the SKILL.md catalog against the request and the recent conversation and verifies the top candidates, so only the picked skill's instructions reach the prompt; the judge is a protocol that a self-hosted model or static rules can implement instead.
- jev-rental
- Consumer rental: sorts every claim in a rental listing into verify-on-site / demand-evidence / high-risk-pitch buckets to build a pre-viewing checklist with code-templated questions; 50-sample calibration reports 0.910 gated accuracy and 0/10 injection flips.
- jev-resume-disqualifier
- Recruiting: knocks a resume out of a pipeline in under 25 ms by asking Jev the disqualifying question first, so only survivors reach a full evaluation.
- Jev-IOT
- Smart Utilities & Telecommunications: Ultra-low-cost, non-autoregressive AI telemetry classifier enabling sub-150ms anomaly triage and autonomic remediation across 10M+ smart meters for under $35/month.
- AgentScope
- Multi-agent platforms: multi-agent platform by Alibaba implementing native TypeSafe Jev classification models for binary, choice, and score routing across agent pipelines.
- inbox-zero
- Email productivity: open-source AI email assistant that uses TypeSafe Jev System One decision models to classify incoming email intent and triage action items.
- SiYuan
- Knowledge management: privacy-first personal knowledge management system featuring native Jev decision model integration for high-speed document classification, flashcard intent categorization, and automated tag routing.
- Paca
- Project management: self-hosted open-source Jira alternative that auto-assigns tasks with a Jev
Choiceover member descriptions, fills blank task fields withChoiceandScorequestions, and routes automation workflows on aChoice/Score/Noulcondition node, applying answers only at 0.6 confidence or above and otherwise leaving the task unassigned or taking the Else branch. - Qualm
- Digital wellbeing: macOS menu bar app that reads the screen as text through the Accessibility API and asks Jev (or Kev, its local open-source counterpart) one
Choiceper user rule plus aNoulon whether the page is a payment, login or banking screen, stepping in with a pop-up only when a rule's probability clears its threshold and never on sensitive pages; on 119 trial pages with Kev, the short-video, feed, livestream and video rules had precision 1.00. - Auto-optimizing Jev: half the errors, 1/7 the cost - Text classification: asks Jev a
Choiceover the readings of a Chinese polyphonic character while the model stays fixed and only the harness around it is optimised, ending at half the errors for a seventh of the cost. - spending-effort-with-jev
- Coding agents: Claude Code plugin whose UserPromptSubmit hook asks Jev a
Choiceover/effortlevels (low / medium / high / max / unclear) plus aNoulon whether a hands-off request has a fuzzy spec, showing a switch tip before Claude starts only at 0.7 confidence or above, with 95% of tips pointing to the right level on a three-rater held-out set. - tab-jev
- Tabular prediction: asks Jev a
Noulon the target plusScorerubrics about each row's text, turns every option's probability into a column next to the row's numeric fields, and lets a tabular foundation model such as TabPFN learn from the labeled rows in context, reaching 0.745 AUC at 256 labels on Kickstarter funding against 0.682 for Jev alone with calibration. - tinystruct-typesafe-sdk
- SDK: TypeSafe Jev integration library for building type-safe classification and routing decisions with structured outputs.
- TetraJev
- General decisions: locally-deployed decision layer for complex decision problems — four readings from two frozen open-weight readers, fused fit-free and routed by agreement with calibrated release gates; benchmarked across eight decision suites plus the RAG reranking pass, including DecisionBench's 35 real-world task categories. No training; does not call the TypeSafe API.
- sortwell
- Personal inbox: MCP server and Claude Code plugin that files each captured note, link or meeting line with one Jev request of
Choicequestions for kind, project and next action plus aNoulfor duplicates, routing to a project only at 0.45 or above and marking a duplicate only at 0.70 with a specific matching item, while the text itself is stored verbatim in append-only local files. - jev-seo
- Site audit: a Claude Code skill that crawls a homepage, has Jev judge page type, search intent, importance, trust and citability, and emits ranked fixes that each carry priority, impact, effort and a source while stating that scores rank work rather than predict rankings.
- IntentSQL
- Natural-language SQL: turns a question about a SQLite database into a sequence of small Jev decisions instead of one generated query, released as an experiment alongside its decision lab.
- TypeSafe Conversation - Home automation: a Home Assistant voice agent built on Jev.
- Jevvie
- Web companions: a page offers its actions as WebMCP tools and one Jev
Choicepicks the action a visitor's request means, with aChoiceper argument asked alongside, asking back when the top two options are close and gating unprompted tips with aNoul(source). - JevRouter
- Coding agents: asks Jev a typed
Choiceover models, subagents, skills, MCP tools, CLIs and plugins, applies availability, permission and risk policies, and returnsno_decisionbelow the configured confidence threshold while preserving the original probabilities. - Gut Check
- Smart home: Home Assistant integration whose eight install checks ask Jev a
Scoreon each pending update's release notes and aChoiceper item elsewhere, such as whether an unavailable entity is expected, worth fixing or safe to remove; answers below 0.5 confidence change nothing, and the rest that need action become Repairs cards the user must confirm. - Vibefilter
- Admin panels: Filament table filter that asks Jev a
Noulper row on a plain-English statement such as "The customer is angry." and keeps the rows at or above 0.8, matching the demo's own mood labels on 861 of 1,000 reviews at about $0.003 and one second per statement, with scores cached by content. - decision-router
- Coding agents: Claude Code plugin, Pi virtual model and CLI that pick the model for each task from one Jev call, a
Choiceover the candidates asked in both orders plus a complexityScorethat sets a minimum tier, raising accuracy on 30 agent-labeled prompts from 73% with theChoicealone to 93%, with quota-aware filtering and a fallback below 0.3 confidence.
Adaptive & Realtime UI
Source file: categories/adaptive-realtime-ui.md
- typesafe-adblock
- Browser tooling: Chrome extension that asks Jev whether each DOM element is an ad, turning ad blocking into a stream of per-element typed questions.
- unclutter
- Browser tooling: WXT extension where Jev decides per page element whether it is clutter, removing it under reusable template rules.
- sift
- Content labelling: Chrome extension that labels every post in an X timeline - substance, humour, chit-chat, promo, junk, or AI-written - with Jev decisions.
- json-render
- Generative UI: Vercel Labs' UI framework uses Jev in its compose path to pick which components and actions a rendered interface should contain.
- PlotVeil
- Spoiler protection: Chrome extension that covers each YouTube comment while one Jev
Noulquestion, batched 20 at a time, answers whether it reveals a concrete plot event of the video being watched or of another title the user protects, with the extension owning the 0.85 / 0.7 / 0.5 threshold and keeping the comment covered when the check fails. - jev-canvas
- Multimodal UI: draw on a tldraw canvas by voice while pointing a webcam-tracked finger; on every partial transcript Jev answers eight typed questions (is it a command, is the sentence complete, action, shape, colour, target, place, size) and plain code gates them with thresholds, in English and Ukrainian, 300–550 ms per decision.
- DWIM
- Desktop productivity: a macOS command palette that reads the frontmost app's menu tree through the accessibility API, asks Jev one
Noulper menu item against the user's plain-language request, and presses the top match when it clears a probability threshold, falling back to a ranked list otherwise and never auto-running destructive items. - SemanticSpace - Semantic mapping: places phrases in 2D by asking Jev how strongly each one relates to two chosen axis concepts and using those scores as coordinates.
- shapeshift
- Input: one text box that morphs into the right UI as you type, asking Jev which control the sentence calls for, and running offline.
- Jevcast
- Desktop productivity: native macOS launcher and window manager that uses Jev to match natural-language window and action commands to known application workflows with local response caching.
Verification & Guardrails
Source file: categories/verification-guardrails.md
- jev-risk-check-provider
- Agent payments: an x402
risk-checkprovider where Jev scores agent counterparties as typed Noul/Choice/Score questions into a code-controlled 0-100 score, issuing an ES256-signed attestation per verdict; 540-call scale run (99.76% at threshold 65-75, 0 false positives) and a 5-iteration 1,500-case adversarial red-team loop (100% adversarial accuracy) with ~$0.00005/decision at p50 ~400ms. - Edward
- Agent operations: one batched Jev
Choiceover the cross-turn coding-agent trajectory decides continue, pause, or escalate, with low-confidence verdicts routed to a human while deterministic code keeps dangerous-command blocking, budget caps, and an Ed25519-signed receipt chain. - is-malicious
- Software supply-chain security: asks Jev
Noulchecks about source and build files, escalates suspicious chunks for a second pass, and returns implicated files and lines before execution. - jev-review
- Software engineering: staged code-review workflow and local dashboard where Jev gates each review stage before a change advances.
- pi-jev
- Agent safety: adds a measured tool-call gate to the Pi coding agent so risky calls are checked by Jev before execution.
- OpenWork
- Engineering workflow: wires Jev into its eval testkit as a verification judge so agent-produced work is gated by typed verdicts rather than a text model.
- jev-guard (leepokai)
- Agent security: prompt-injection and dangerous-action guard for Claude Code, Codex, Pi, and ACP agents, with Jev deciding what to block.
- Foreman
- Software factory: sits above Codex workers and has Jev independently judge whether an implementation is complete, its tests sufficient, or a human is needed.
- stanley-code
- Coding agents: bounded Jev workflows that keep agent judgments typed instead of free-form.
- opencompany
- Agent workspace: runs its approval review through Jev so workspace actions are gated by a typed decision.
- jev-git
- Developer tooling: sub-second Git pre-commit & pre-push reflex gate that screens staged diffs for secrets and destructive commands using Jev.
- pi-heed
- Runtime constraints: checks every side-effecting tool call from the Pi agent against what the user actually asked for.
- Hunch (Kelbie)
- Code review: plain-English rules that Jev checks code against, locally or on every pull request, with Jev picking one label per finding.
- Abide
- Agent supervision: reads every edit a coding agent makes and has Jev flag rule violations, with the project reporting that an independent reviewer confirmed 10 of the 39 flagged edits and 11 of the 15 flagged turns.
- fx
- Coding agent: ships a
typesafe_permission_reviewerbuiltin so the agent's permission decisions run through Jev rather than an LLM call. - Sniff Test
- Writing: prose linter that asks Jev ten
Booleanquestions per paragraph (stacked hedges, restating closers, not-X-but-Y turns, naked cost figures) at a 0.7 threshold; CLI, pre-commit hook, GitHub Action and Claude Code skill; measured 182 ms median and 1 of 54 clean paragraphs flagged against 37 for Haiku 4.5. - jev-pref
- Code review: turns the preferences in a project's AGENTS.md into
jev-pref.jsonrules that Jev checks against each diff hunk, staged file set, or pull request, returningfix_nowor advisory findings to the coding agent and a nonzero exit code on blocking ones. - jev-axi
- Agent safety: PreToolUse gate for Claude Code and Codex that has Jev score each shell command for destructiveness, exfiltration, remote code execution, and security weakening, deciding routine commands locally so nothing is sent for them, and scoring 44/44 on the 44 labeled tool calls in its repository.
- pi-verdict
- Agent safety: Pi permission gate where Jev answers one Choice (allow/ask/deny) per gray-zone tool call — deterministic rules settle clear cases first, deny blocks, ask escalates to a human confirm, and errors or timeouts deny; Jev is an optional backend, experimental, reached through OpenRouter or TypeSafe's direct API.
- jev-commit
- Developer tooling: pre-commit hook where one Jev call judges whether the commit message matches the staged diff, flags debug leftovers and unmentioned work, and blocks only on a detected credential.
- Blink - Code review: CLI that coding agents run after every change, with Jev checking the diff near-instantly in place of an LLM reviewer.
- hermes-jev-approvals
- Agent approvals: proof of concept that puts Jev in front of Hermes Agent's command approvals, reporting 8.7x faster decisions and 4.4x fewer prompts to the user.
- taste-lint
- Writing / UI: CLI that uses Jev probabilities on semantic taste checks to catch AI slop in UI, copy, and agent instructions before ship; measurable rules stay local and active findings can fail a run.
- jev-engineering
- Agent safety: gates coding-agent tool calls with deterministic rules first and one typed Jev call second, then publishes a rerunnable 300-call injection test showing what the gate catches and what walks past it.
- jev-harness
- Developer tooling: System 1.5 quality gate and token optimizer for AI coding agents that triages test failures in < 500 µs to resolve missing dependencies without frontier LLMs, aborts circular doom loops, and modulates reasoning effort across Python, TypeScript, and Rust.
- Reflex
- Coding agents: Pi-based coding agent that sends each state-changing tool call through one Jev request of five
Noulrisk checks plus a riskScore, maps the answers in code to allow, ask or block by the user's risk setting (protected paths always ask), and also uses Jev to pick the model tier per prompt and to send back "done" claims that ran no verification, at about 400 ms per decision. - r2r-jev
- Agent governance: asks Jev two
Noulchecks per tool call (beyond scope, destructive) and admits each judgment as Evidence that can degrade Trust, Delegation, and Authorization until a human override repairs the relation, so later calls inherit the history; includes a stateless-vs-stateful comparison with a scenario adversarial to persistence. - GeekLink Jev Subtitle Translator
- Subtitle translation: asks Jev a
Noulreview question for each translated subtitle line to flag omissions, changed meaning, names, numbers, negation, or other defects for human review before export. - TryJevAI - Scheduling: public Jev playground uses a typed
Choicewith an explicitUnresolvedoption to distinguish a mentioned arrival time from an agreed meeting time, showing the returned probabilities and prompting for missing agreement before treating a time as settled. - Agent Chaperone
- Agent safety: MCP proxy plus hooks that screen a tool call before it runs and a tool result before the agent reads it, with 45 test files behind it.
- jev-proof
- Creator sponsorship: verifies each sponsored ad segment in video subtitles against acceptance rules with one Jev
Noul+Choicecall per fact while deterministic code keeps the confidence gate; 90-sample calibration reports 0.922 gated accuracy and 0/15 injection flips. - jev-fidelity
- Editorial QA: asks Jev per fact unit whether an edit preserved the original (preserved / equivalent / drift / lost) behind a 0.70 confidence gate in code, degrading to human review rather than pass; 55-sample calibration on real Wikipedia revision diffs reports 91/92 gated judgments correct and 0/20 injection flips.
- approval-judge-bridge
- Agent safety: OpenAI-compatible /v1/chat/completions proxy that gates an agent's shell commands through a calibrated Jev Choice decision with fail-closed semantics.
- Dub
- Link safety: calls
typesafe-ai/jevinmalicious-link-check.tsbefore a short link is created, so the URL is gated by a typed verdict rather than a blocklist. - Canny
- Agent verification: stops AI coding agents from claiming work is done without evidence by using deterministic hooks and TypeSafe's Jev advisor to evaluate test results, file diffs, and verification logs.
- JevGate
- Code review: CI and coding-agent gate that parses code locally and asks Jev
Noul,ChoiceandScorequestions about one function, file outline, candidate copy pair or test at a time, turns answers at 0.80 intorevieworconsiderfindings with file and line, fails the build onreview, and keeps undecided files asuncertaininstead of clearing them. - dsh-jev-interceptor
- Coding agents: DeepSeek Harness plugin where a Jev
Choicerisk class plusNoulirreversibility, task-match, and injection checks gate every non-read-only tool call (deny confident high-risk, ask ambiguous, delegate the rest),Noulscope and reversibility questions auto-approve clearly-granted calls behind argument-evidence gating, and a per-messageScorere-ranks what a referenced session keeps instead of oldest-first dropping — fail-closed to stock behavior, shadow mode with a/jev-statscommand, 64 tests. - claude-code-templates
- Agent safety: CLI configuration suite for Claude Code featuring a
jev-guardrailsmod that screens prompts and turns against jailbreaks, harm, and policy breaches via TypeSafe System One. - jevci
- Quality gate: asks four typed questions about each change — three
Scorelenses and oneNoul— and blocks a diff, commit message or doc set that falls below the resulting quality score, from the terminal, a pre-commit hook or a GitHub Action. - pi-subagent-jev
- Agent governance: when the Pi main agent dispatches a subagent, evaluates the task text against configurable rule sets in one typed Jev call (per-rule probability questions with below/above thresholds), blocks the dispatch with per-rule reasons on any hit, and fails open to allow on errors.
- jev-lint
- Software engineering: uses Jev
Nouljudgments and local thresholds to flag team-rule violations as Claude Code and Codex edit, helping agents fix them before code review with configurable rule packs and repository-specific rules. - jev-secret-guard
- Agent security: Claude Code PreToolUse hook that blocks known key formats locally and sends unknown high-entropy strings to Jev only in masked form for a
Noulon whether they are real credentials, blocking at 0.80 and asking the human from 0.30 or whenever Jev is unavailable; 6 of 6 secrets and 0 of 6 benign strings were blocked in its published calibration. - Perch
- Code linting: semantic code linter that asks Jev about each method with its callers and callees in view, a
Noulfor whether it has a bug, aChoicefor which kind and which line, and aScorefor severity, plus language-filtered CWENoulchecks and custom rules written as sentences at repository, file or method level, failing CI on any answer over its floor. - semcheck
- Code review: Go linter whose rules are plain-English questions such as "does this log call write personal data?", asking Jev one
Noulfor each piece of code a rule applies to and reporting it above the rule's threshold; its two shipped rules were right on 12 of 12 sampled findings in three open-source projects. - Cribrix
- Retrieval / RAG: filters retrieved chunks with a Jev
ScoreplusNoulchecks for answer evidence and prompt injection, then withholds any draft whose claims fail a batched per-claimNoulor cite numbers absent from the sources; on its replayed 62-question golden set it answered 0 of 22 unanswerable questions, against 4 of 22 for naive top-5 RAG. - Skill Scanner
- Agent security: Cisco's scanner hunts prompt injection and exfiltration in agent skills, and ships a System One analyzer as a deliberately advisory tier that cannot emit a finding or change a severity.
- openclaw-jev-leakguard
- Agent security: OpenClaw plugin that checks every outgoing agent message against where it is going, running local key-format and term rules and then five Jev
Noulquestions in one call (credential, where credentials are kept, client name, internal infrastructure, confidential business information) through OpenClaw'sdecisionModel, hosted Jev or a local Kev, and blocking, asking or holding it back by the channel's public, shared or private tier; with Jev it missed 0 of 56 synthetic leaks, 30 of which no regex or term list could see, with 4 false alarms on 57 ordinary messages at 223 ms p50.
Scoring & Ranking
Source file: categories/scoring-ranking.md
- Clean Code Judge
- Code quality: scores every file of a pull request on 31 boolean Clean Code smells plus function size and nesting, then hands the verdicts to a writing model for the review prose.
- citation-verifier
- Academic publishing: checks whether each cited paper actually supports the sentence citing it, with Claude locating the quote, Jev scoring the support, and a human making the final call.
- jev-assist
- Coding agents: ranks every tracked file by relevance to a one-line task description — Jev asks each file the same typed question in parallel batches, so an agent in a 600-file repo starts from the handful it actually needs — with a validate command that grades the ranking against past commits.
- jev-ai-detector
- Writing analysis: Chrome extension which gives readers an instant, uncertainty-aware signal for how strongly selected webpage text resembles AI-generated writing, using Jev inline in Chrome without interrupting reading.
- jev-bfs
- Search tooling: finds link paths between English Wikipedia articles by having Jev rank each page's outgoing links while Python controls the search.
- Jev Search
- Web search: uses Jev Noul judgments on result titles and snippets to rank Search1API results by relevance, with application code merging duplicate URLs and grouping lower-scoring matches separately.
- Tweet Radar
- Social reading: uses Jev
Noulto score already-loaded X posts against a reader's goal and profile, then pairwiseChoicejudgments to rank eligible matches and surface up to three for review. - pagegrade
- Content quality: grades page sections for clarity, writing, and on-page SEO with Jev and returns per-section scores.
- jev-scout
- Developer tooling: sub-second zero-hallucination open-source repo and crate scout using TypeSafe Jev speculative fan-out scoring.
- jev-seo
- Zero-cost, agent-first SEO & Generative Engine Optimization (GEO) search radar CLI suite and MCP server powered by DuckDuckGo and TypeSafe Jev System One.
- JevSlop
- Writing quality: scores public note.com articles on eight Jev
Scoreaxes inside a singlesystemOnerequest and turns them into a 0-100 Slop Score in ordinary TypeScript. - Supercov
- Code quality for coding agents: Jev answers twelve
Noulproperties per source file so the agent knows what to fix first. - jev.nvim
- Developer tooling: Neovim plugin that splits the buffer into functions with Treesitter, scores each against a plain-language question with Jev, and ranks answers by probability in quickfix.
- Jev RAG
- Retrieval and RAG: shortlists local document passages with BM25, reranks up to 30 candidates through one batched set of Jev
Noulrelevance questions with an optional threshold, and passes the strongest evidence with source citations to a configurable answer model, backed by a runnable local demo and a public 323-query NFCorpus benchmark. - jev-reranker
- Retrieval and RAG: uses Jev Noul judgments to assess retrieved documents for relevance and usefulness as answer evidence, then sorts results and optionally filters them using a configurable threshold.
- Jev Reranker (Rust CLI)
- Retrieval and RAG: JSON-in/JSON-out CLI that uses separate Jev
Noulchecks to rank candidates, apply evidence thresholds, or extract source text while keeping those decisions independent. - jev-skip
- Media: browser extension that reads the YouTube caption track and scores each segment's sponsor probability on the seek bar before the intro ends, reporting 77% of SponsorBlock's sponsor seconds caught over 23 videos at $0.0008 a video.
- jev-semgrep
- Semantic search: greps by meaning across languages, having Jev score every line against a meaning and letting meanings combine with AND, backed by a 13-file test suite.
- nlgrep
- Developer tooling: uses Jev
Nouljudgments to find code, docs, logs, and text satisfying natural-language conditions, with a configurable probability threshold and ranked file results linked to source lines. - JevPDF
- Document search: in-browser PDF viewer that extracts each page's lines locally with pdf.js and asks Jev one
Noulper line on whether it answers the query (16 lines per request, sharing the page text as state), highlighting lines at or above 0.55 page by page and ranking them by probability. - slop-grader
- Content quality: CLI tool that grades text files against custom rulesets for AI slop, grammar, and technical doc quality using Jev scores and line-level flags, then guides an AI agent to auto-fix violations.
- jselect
- Research and retrieval: selects source-linked evidence within a token budget using Jev Noul relevance judgments and local diversity-aware selection.
- Jev Deep Research
- Evidence retrieval: uses Jev
Choiceto locate source lines andNoulto check evidence presence across document regions in parallel, then returns original passages to a GPT research agent through Pi-Serini with 20/40/60-document batch limits. - jsort
- Text measurement: ranks text along a plain-English criterion using pairwise Jev Noul comparisons and a locally fitted Bradley-Terry scale.
- jgrep (kyu1204)
- Developer tools: semantic grep that asks Jev one Noul per 5-60 line code chunk, diff hunk or CSV row (16 per request) and prints grep-style file:line hits above a threshold, so English sentences work as CI lint rules.
- jev-resume-screening
- Recruiting: screens one resume against a JD in a single request of five Noul evidence gates, four Score dimensions, and one background-routing Choice, with criteria hardened v1→v3 against negative-control resumes (a glossy-trap CV's self-described "AI heavy user" fell 0.95→0.49) and any low-confidence answer escalated to human review.
- hippo-memory
- Agent memory: a biologically-inspired memory store whose optional Jev reranker lifts recall R@1 from 0.41 to 0.62 on a private 300-query developer store.
- MemSearch Jev reranking
- Coding-agent memory: an optional Jev reranker asks Noul questions about retrieved Markdown chunks and sorts them by relevance to the query, with bilingual evaluation results.
- Oko
- Developer tooling: local code search for coding agents that shortlists function-level chunks with ripgrep and BM25, asks Jev a
Noulrelevance question per chunk across three parallel requests, and returns the accepted ones as excerpts through MCP; the cutoff and excerpt selection live in code. - grokbot-jev-jobs
- Job search: a daily Vercel cron that scores public job postings against one resume with Jev through the Vercel AI Gateway, so only the plausible matches surface.
- jeff
- Developer tooling: Go CLI whose rank command asks one Jev
Scoreper item per weighted dimension of a YAML spec in a single request and sums weight times score in code to order the items, with noul, choice and score commands that turn a threshold into exit code 10 for shell scripts and CI. - Paper Radar
- Research: scores every new arXiv and bioRxiv paper against plain-English interests with one Noul each and publishes the top picks as a daily page and RSS feed.
- Refix - Growth: asks Jev a
Scoreover each experiment result to decide whether it clears the promotion bar, and aChoiceover candidate plays to decide what to run next in SEO, content, and ads. - OpenViking
- Reranking: Volcengine's agent context database ships a Jev rerank client that scores each candidate document with
jev-latestagainstapi.typesafe.aiand treats the returned probability as relevance, because TypeSafe exposes no native rerank endpoint. - jevsearch
- Site search: shadcn/ui command-palette block that streams keyword hits on the first keystroke, then sends the top 20 to Jev in one request (a
Noulper page on whether the visitor would be glad to land there, aChoicefor the single best answer, and aNoulon whether any page answers at all) and re-orders or drops hits in code, with the repo's own benchmark over the 109-page TypeSafe docs reporting Hit@1 of 83% against 41% for its keyword pass alone. - jev-retrieval
- Coding agents: Rust CLI (
jevr) that turns a natural-language query into grep-stylepath:start-endtargets — a stateless local BM25 pass proposes candidates, JevNoulmembership questions score their 100/20-line windows (kept at 0.90 for code, 0.60 for docs), and one listwiseChoiceper lane orders the keepers — ships as a Claude Code skill and plugin, and placed 2nd of 90 models on the HAKARI-Bench NanoRTEB reranking leaderboard. - Vector Graph RAG
- Multi-hop retrieval: uses Jev Noul judgments to score candidate relations and applies a configurable threshold before retrieving their linked documents.
- Jev-Code-Reviewer
- Code review: asks Jev for a priority score per changed unit and returns a
priorityGapthat a local uncertainty policy turns into the order a human should read the hunks in, while OpenAI explains the ones that surface. - WorldMonitor
- Geopolitical intelligence: real-time global intelligence dashboard using TypeSafe Jev questions to score news headline severity into 5 threat levels and categorize events across 14 conflict, cyber, and infrastructure domains.
- LinkScout
- Web search: browser extension that asks Jev, in one request per result, for relevance and depth
Scores,Nouls on SEO filler and sales pages, a page-typeChoiceand aChoiceover pre-split paragraphs for the key passage, then combines them in code into a 0–100 badge that re-ranks Google, Bing and DuckDuckGo results.
Agent Decisions
Source file: categories/agent-decisions.md
- Learn Jev end to end
- Developer education: a 12-notebook Python course whose hand-rolled agent loop asks Jev a
Choice(allow / ask / block) withNoulirreversibility and exfiltration checks before every tool call, sends ask verdicts to a human and fails closed on errors, and adds aChoicemodel router with a confidence fallback and aNoul"am I done?" gate, each measured against labeled fixtures. - Hermes JIT Context OS
- Coding agents: uses Jev as a sub-millisecond System 1 Epistemic Gate and Domain Router to score AST relevance, test proofs, and tool targets, cutting autonomous agent turns by 31.3% and blind file exploration by 52.6% on SWE-bench with fail-open circuit-breaker resilience.
- Jev by Example
- Agent development: runnable JavaScript lessons use Jev Choice, Score, and Noul judgments for memory reconciliation, recovery proposals, and handoff checks, with explicit application policies, offline fixtures, and opt-in live calls.
- jev-social
- Social media research: uses a Jev
Choiceat each step to select a concrete socai CLI operation and observed post or profile target on Instagram, TikTok, or LinkedIn, rejecting malformed or low-confidence decisions before execution. - Jev Ultrafast
- Browser automation: browser-use's ultrafast agent where Jev decides each next action and element to click, calling a language model only when text must be typed.
- jev-agent-browser
- Browser agents: a parent agent delegates bounded tasks to a Jev loop that selects typed browser actions, validates them through agent-browser, and escalates ambiguity or stuck states back to the parent.
- pi-typesafe-jev
- Coding agents: exposes System One judgments as five Pi tools so a model makes narrow semantic judgments while code and users keep control of thresholds, weights, and actions.
- jev-judgment
- Coding agents: agent skill that sends closed coding-agent judgments to Jev so verdicts stay typed, cheap, and comparable across runs.
- limpet
- Coding agents: Stop hook that keeps an agent from finishing too early by judging plain-language completion rules with Jev.
- dsh-auto-mode
- Coding agents: DeepSeek Harness permission preset whose end-prompt step has Jev answer the open questions an agent leaves in its final message, steering them back only when a choice clears 0.6 confidence and an autonomy-safety Noul clears 0.5, and returning the turn to the human otherwise.
- augustus
- Coding agents: independent augustus and augustus-train skills for application-specific decision models, covering primitive/base-model/method selection, data assembly, fitting, export/reload, bounded improvement and independent evaluation, with TypeSafe Jev as the default hosted exemplar.
- yoshi
- Context management: proxy for Claude Code and Codex where Jev judges which conversation history is still needed before pruning.
- pi-jev (TheoOliveira)
- Coding agents: semantic tool routing and typed System One decisions for the Pi coding agent.
- pi-quiet-ask
- Coding agents: gives the Pi agent a quiet Jev decision layer for judgments it would otherwise hand to a chat model.
- fastbrowse
- Browser agents: Jev picks each action from what is on the page while an LLM reads and plans.
- super-jev
- Decision harness: turns a Jev answer into a bounded action instead of leaving the caller to interpret it.
- jev-superpowers
- Coding agents: software development framework for AI coding agents that hands package vetting and completion gates to Jev typed decisions.
- Jev Browser
- Browser automation: drives a browser with Jev deciding each step, pitched as fast and very cheap next to LLM-driven browsing.
- pi-fast-jev-compaction
- Context management: Pi extension that keeps conversation text verbatim while pruning stale tool history with Jev, falling back to Pi's own summarization only when pruning cannot free enough room.
- Atomic
- Coding agent runtime: ships a first-class Jev structured-output provider so an agent's decisions come back typed, through the same decision resolver as its other providers.
- fast-jev-compaction
- Context management: Claude Code plugin that replaces the compaction summary with Jev decisions, scoring every tool call and result for whether it is still needed instead of summarizing the session.
- fast-dev-compaction
- Context management: Codex port of the Jev-guided compaction idea, restoring context verbatim around a session compaction rather than summarizing it.
- public-browser
- Browser control: lets Claude Code and Cursor drive a real Chrome profile, with a Jev loop deciding the actions, reporting roughly 30% fewer tokens and 25% lower cost.
- pi-typesafe-router
- Coding agents: routes Pi's work through typed Jev decisions.
- wakegate
- Long-running agents: before a sleeping agent's LLM is resumed on a timer or incoming event, Jev answers a
Choice(wake, not yet, unrelated) against the agent's own sleep note, and code skips the wakeup only when wake is below 0.2 while always waking on user messages, bare timers, a skip limit, errors, and timeouts; one run passed 21 of 21 hand-written scenarios, which the README calls a smoke test rather than a benchmark. - BrowserClaw
- Browser automation: Zero-lock, session-preserving Chrome MCP server that couples a local Jev System One semantic micro-loop (
chrome_act_toward_goal) with an 85%+ pruned DOM tree (Shadow DOM & iframe pierced), dispatching native CDP events (isTrusted: true) on active logged-in sessions without focus theft. - jev-belay
- Coding agents: Claude Code Stop hook that reads the transcript for evidence and spends one four-question Jev call only when files changed with no passing check since, failing open on any error.
- Jev for Chrome
- Browser automation: unofficial Chrome extension port of Jev Ultrafast where a Jev
Choicepicks the operation and DOM element each step and twoNoulchecks (goal reached, stuck) veto a premature DONE or BLOCKED, with a small text model used only when text must be typed. - jev-pruner
- Context management: Claude Code plugin that trims long Bash output with Jev before the model ever sees it, keeping terminal noise out of the window.
- jev-desktop
- Computer use: supplies Jev action selection inside Codex Computer Use, choosing among desktop actions rather than asking a language model at every step.
- jev-agent-skill
- Developer tooling: Claude Code/ZCode skill that offloads classify/route, batch-screen, score, and compliance-check judgments to Jev via OpenCode Zen's free tier, bundling a zero-dependency jev.py caller (transient-500 retry, WAF-safe UA, GBK-pipe-safe stdin) and a production Taobao-shop comment-triage pipeline that keeps raw items out of the agent context.
- Yappy - Computer use: macOS voice agent that asks Jev one
Choiceper step (operation and target control) over the front window's accessibility table, executes only validated high-confidence answers, and escalates to a full LLM agent on low confidence, no-effect actions, or unknown field values; author-reported 275–690 ms per decision. - JevLoop (zjunlp)
- Agent harness: routes the loop's own judgements to Jev, where a
Choicepicks the next tool from candidates rebuilt every step, aScoregrades the call's risk, and aNouldecides whether it needs authorisation, while plain code acts on the answers so a high risk score forces human authorisation that no probability can override (7.7% of wall clock with the offline judge, 79% over the hosted API). - JevLoop (parkavenue9639)
- Agent runtimes: a Python runtime where Jev
Choicedecisions select tools and targets, uncertain decisions escalate to an LLM, and a shared guarded kernel supports isolated Docker workspaces and paired LLM-only comparisons. - DataJev
- Data analysis agents: an LLM performs Python-based analysis while Jev reads the compressed analytical state and decides whether the agent should continue the current direction, switch to another one, verify a finding, or stop and synthesize the answer.
- jev-mobile
- Mobile control: fast structured Android control loops that route each step through Jev alongside Mobile MCP, with 35 test files.
- GUI JEV Harness
- Computer use: recursive screenshot grounding where Jev returns a
Choiceover grid-tile candidates at each level, and local probability and margin gates decide whether to descend or refuse, emitting only a raster point and bounding box and never clicking. - jev-compaction
- Context management: standalone agent context compactor where Jev only scores transcript segments — kept lines stay verbatim, low scorers move to a store behind an expand() pointer instead of being deleted, and the append-only frozen prefix keeps the prompt cache valid; runnable offline demo, no API key needed.
- Visual-JEV
- Multimodal models: Jev-style model built on Qwen3.5-4B that takes images directly, without first converting them to text.
- DeepSearcher stopping-policy experiment
- Agentic search: a standalone evaluation uses Jev Noul judgments on accumulated evidence to decide whether to stop or continue within a search-round budget, comparing stopping behavior, evidence recall, and decision cost.
- neo4jev
- Graph navigation: navigates a Neo4j knowledge graph hop-by-hop using Jev Choice over candidate outgoing relationships and Noul to detect goal completion, using beam search over answer log-probabilities.
- jev-chat
- Messaging: an Android accessibility service reads the conversation in WeChat, QQ, X, or Feishu, asks Jev
Choiceover candidate replies, and fills the draft box while sending stays manual; a Windows port does the same from offline OCR of the WeChat window. - hermes-jev-skills
- Agent runtime: a Jev-powered skill suite that decides model routing, memory, compaction, skill selection, and computer or browser use for Hermes agents, and also installs under Claude Code and Codex.
- jev-browser-use
- Browser automation: lets Jev pick the click while Codex thinks and verifies, reporting 5-10x faster browser operations behind four CI-run contract tests on the bridge.
- mobile-jev
- Mobile agents: puts Jev into on-device screen-aware action selection for the Droidrun loop, with a Jev Studio web app streaming live device and decision telemetry.
- Jev-cu
- Computer use: drives a GUI through Jev decisions with a
jev-decidescript and ships a P0 case set taken from accessibility-tree snapshots of a calculator, a calendar, and the NetEase home screen. - SkillRanker
- Coding agents: standalone Rust CLI that uses Jev to rank candidate skills against live session context, advising the next step through a Claude Code UserPromptSubmit hook.
- AutoGPT
- Autonomous agents: open-source autonomous agent platform featuring first-class TypeSafe Jev decision blocks for typed routing, filtering, scoring, and confidence-gated next-action dispatching.
- dsh-jev-decide
- Coding agents: DeepSeek Harness plugin whose single
jev_decidetool lets the agent ask aNoul,Choice, orScorequestion about any state — urgency triage, intent routing, guardrail checks — and gate on the returned probability or confidence in code instead of trusting the chat model's guess. - jev-browser-bridge
- Browser agents: plugs any CDP browser into a Jev loop, where a Jev
Choicepicks the operation and its target element each step from candidates read off the DOM rather than the layout, so the same agent runs on Chrome and on engines that never draw a page (Moli, Lightpanda, Kitesurf), passing at least 90% of runs on each of fourteen browsers tested. - Eliza
- Autonomous agents: multi-agent framework integrating TypeSafe System One decision services for sub-100ms intent classification, action dispatching, and confidence-gated tool execution.
- oh-my-claudecode
- Coding agents: multi-agent team orchestration for Claude Code featuring opt-in Jev hooks for sub-millisecond judgment points, decision caching, and per-point egress controls.
- jcode
- Agent runtimes: RAM-efficient autonomous agent harness implemented in Rust with native TypeSafe Jev typed decision transport for memory pruning, browser navigation, and voice interaction routing.
- opencode-jev-compaction
- Replaces OpenCode compaction summaries with Jev keep/drop judgments that prune stale tool calls while preserving everything kept verbatim.
- jev-opus
- Coding agents: runs Claude Code on Opus 5.5 and asks Jev a
Choice, aScoreand aNoulon each prompt and after every tool batch to pick the next API call's reasoning effort (low/medium/high), sent as a per-message statement so the prompt cache never breaks. - jev-auto-approve
- Coding agents: Claude Code PreToolUse hook that asks Jev a
Noulon whether a shell command is strictly read-only, auto-approving at 0.95 and otherwise falling back to the normal permission prompt without ever denying, while a local hard-no list and injection filter keep risky commands from reaching Jev; 0 of 8 state-changing commands were approved in its published calibration. - WebJev
- Browser agents: open-weight Apache-2.0 decision model (a Qwen3.5-35B-A3B fine-tune) that answers jev-ultrafast's per-step
Choicequestions for the next operation and its target element behind the same/v1/systemoneAPI, completing 38.5% of 125 hand-picked real-website tasks graded by deterministic verifiers versus 16.7% for Jev 1.13 in the same agent. - laya-browser-agent
- Browser agent: derives each step from a Jev-shaped model — Laya through MLX or PyTorch, any duck-typed backend, or an arbitrary System One HTTP endpoint.
- openclaw-jev-trigger
- Agent automation: OpenClaw plugin and CLI that turn a plain-language
--when/--not-whencondition into a scheduled trigger script, asking Jev oneNoulper tick through OpenClaw'sdecisionModeland waking the conversation model only when the condition becomes true at 0.7 or above; on 76 synthetic watcher ticks Jev was right on 75 with 0 false wake-ups at 231 ms p50 and about $0.000016 per check, against 87% for first-try JavaScript rules. - Sedum
- Browser end-to-end testing: in goal mode each turn asks one Jev
Choicefor the next operation (click, type, done or blocked) plus a speculative target among the page's offered elements, capped at 24 requests, 18 actions and 120 s, and the test passes only when an independent verify claim clears twoNouls (holds ≥ 0.75, contradicted flagged at ≥ 0.5), since the planner's done is never a verdict; authored-step tests reuse the sameChoiceto resolve each plain-English step, and with your own API key a 20-person team's PR suite costs $38–$91 a month vs $4,875 on a per-step AI platform.
Data Labeling & Curation
Source file: categories/data-labeling-curation.md
- jev-align (Sutro)
- Dataset engineering: evaluates CSV, Parquet, and JSONL rows with Jev
Choice,Score, orBooleandecisions, sends ambiguous and audit samples to a human, and uses accepted human labels to optimize the saved definition with GEPA. - jev-curate
- Dataset engineering: sifts synthetic JSONL and Parquet rows using Jev Noul checks and calibrated confidence scores, streaming passed records and rejections straight to disk.
- typeful-triage
- Open-source maintenance: multiplayer triage dashboard where Jev answers a fixed set of typed questions per issue — kind, severity, urgency, duplicate, and next step — and every human correction is kept and shown back to the model on later runs.
- jlink
- Research data: links records under a plain-English match rule using Jev Noul pair judgments, with local candidate blocking and match resolution.
- jgrep
- Data filtering: filters text, structured records, functions, and diff hunks against plain-English descriptions using Jev Noul judgments.
- jevgrep (allebee)
- Log triage: filters logs and other text streams, including live
tail -foutput, by asking Jev one Noul per line against a plain-English question and printing lines at or above a probability threshold, with a hand-labelled benchmark against Claude in the repository. - jev-research-pipeline
- Research monitoring: asks Jev Noul gates and Score dimensions per (paper, research question) on each daily fetch through Pydantic AI's typesafe model, keeps sources above a code-side threshold, and hands them to Qwen for question-centric Obsidian notes; offline tests replay recorded cassettes.
- GroundingJev
- Visual annotation: a Jev-inspired Qwen3.5-0.8B model that maps an image and referring expression to four bounding-box coordinates in one forward pass, reporting an 8.61× inference speedup over its autoregressive base model.
- jevextract
- Information extraction: LangExtract alternative where code proposes candidate spans with exact offsets and Jev answers one
Choiceper span (a schema class or none) plus aNoulper sentence-level class, keeping answers above a per-class threshold and flagging close calls for review, with a published benchmark measuring 10–26× lower cost than LangExtract on Gemini 3.5 Flash but lower F1 (84.2 vs 88.5 on its bilingual jx-bench). - JevSpan
- Information extraction: zero-shot named entity recognition that splits text at punctuation, asks Jev one
Choiceover every candidate window per entity type, verifies each nominee with a secondChoice(the type, none, mixed or partial) and settles its boundary with a third, averaging 73.7 strict F1 across 12 Chinese and English NER benchmarks against 72.1 for direct extraction with Qwen3.8-27B.
Evaluation & Benchmarking
Source file: categories/evaluation-benchmarking.md
- Jev Web Analyzer
- Product evaluation: analyzes a public SaaS landing page as clean Markdown and asks Jev ten bounded
Choicequestions about first-visit understanding, returning inspectable findings for the first change to make. - Jev Playground
- Model evaluation: benchmarks Jev against Luna, Haiku, and Gemini at choosing validated legal moves in explicit-state games, scoring decision quality and consistency across a sequence of moves.
- Jev vs Mistral and Gemini for event validation - Event discovery: head-to-head test of Jev against Mistral Small and Gemini Flash-Lite at validating local event listings.
- jev-research-eval
- Research automation: reproducible eval harness plus field note for Jev Ultrafast research-browser tasks, with QC'd cases, a suite runner, and a report generator.
- Jev judge call vs dimension scores - Model evaluation: tests one direct Jev question per row against 12–14 Jev-scored dimensions with locally fitted weights on three classification tasks, reaching 0.9076 against 0.8373 on Japanese NLI but flagging about 25× more hard benign rows as attacks.
- Jev Pong
- Model comparison: Pong where the ball advances one step per model decision, putting Jev head-to-head with LLMs through Vercel AI Gateway.
- Jev reranking is not a free win - Search reranking: a measured run over 33,047 catalog entries, 164 real queries, and 9,831 graded pairs reports that Jev reranking alone did not beat vector retrieval.
- An early-access test of TypeSafe's Jev - Independent trial: measures calibrated judgments on early-access Jev and reports the resulting cost per decision.
- jevcal
- Model evaluation: fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, reports how much traffic still has to escalate to an LLM, and fails CI when a model update breaks the locked thresholds.
- WindTunnel
- Browser-agent benchmark: measures WebMCP against other browser-agent interfaces, with Jev appearing as one of the compared configurations.
- jev-eval
- Third-party check: compares Jev against GPT-4o-mini and Claude Sonnet 4.5 under identical conditions on the same judgment task.
- minutes
- Meeting notes: local-first transcription app whose live voice path runs its evaluations through Jev.
- jev-orderby-bench
- Model evaluation: measures whether a SQL ORDER BY over a Jev probability is defensible (pairwise inversion, Score ordinality against a human grade, calibration, wording invariants, sort-key ties) under a pre-registered gate that jev-1.13.0 passes on 20 Newsgroups topics and fails four of six conditions on Amazon ESCI product relevance, and shows a DuckDB extension's default 40-row batching fails the ranking gate that one row per request passes.
- jev-ood-calibration
- Model evaluation: independent calibration test of Jev on 900 rule-generated support tickets it cannot have seen plus three public benchmarks, publishing every raw response, ECE against a simulated noise floor, temperature refit, and the per-type sign of miscalibration (Choice and Score overconfident, Boolean underconfident).
- ASSAY-001
- Independent pre-registered check of Jev calibration and type safety on Banking77 / CLINC150, with a split verdict and full logs, written up at donttrustme.ai.
- BTK audit studies - Content & growth: Jev striking-distance triage ranks SEO fixes and drives study pages; 1,204 pages judged per run, 4,816 judgments in under 3 minutes, $0.0048 per 12-query batch.
- Can Jev Be a Better Agent Evaluator? - Agent evaluation: LangChain compares Jev against LLM judges on accuracy, repeatability, latency and cost, concluding Jev is the cheaper and more consistent judge for online evals.
- jev-acento
- Language evaluation: pre-registered paired audit of Jev on Spanish over 3,200 human-labelled items, finding that a Spanish
statecosts 3.0-6.4 pp of accuracy and roughly doubles ECE on XNLI and PAWS-X while writinginstructionsin Spanish changes nothing, and shipping a CLI to rerun the same comparison on your own labelled data. - Jev vs GPT-4.1 on a synthetic survey ![stars](https://img