← 开源
QingYunA

answer-me-with-htmlNEW

Answer me with HTML — an agent skill that answers hard questions with a one-page HTML you can actually read. 让 AI Agent 用一页 HTML 回答复杂问题。

AI EngineeringGive agent toolsPrompt & ScaffoldJavaScript
在 GitHub 打开
增长势头
+47824 小时新增 Star+60.0%
1.27k
Star
86
Fork
—
本周
7
贡献者
创建于 2026-10-02 · 更新于 2026-10-05 · 今日第 25 名
主要开发者
README

Answer me with HTML logo

Answer me with HTML

An agent skill. Ask a hard question, get a page you can actually read instead of a wall of text.
The model writes about 1/7 of the tokens it would need to hand-write the HTML.

License: MIT CI Works with Claude Code, Codex, Cursor, OpenCode

English · 简体中文

Once installed, ask questions the way you always do:

> Explain the TCP three-way handshake
> Map out how the modules in this repo fit together
> Redis or Memcached for our cache?

The agent writes a short Markdown draft and hands it to the CLI that ships with the skill. About 50 ms later you have a page:

https://github.com/user-attachments/assets/d3063a28-5dfd-4c44-a562-be901c49b249

24-second demo. Turn the sound on for the music.

Why not just ask for HTML?

You can. Models write decent HTML now. The cost is output tokens: the model has to type every line of CSS, every wrapper div and every SVG coordinate, and output tokens are what you sit and wait for.

With this skill, the model writes only the content. We asked the same questions with the same model both ways, in a plain Claude Code setup (3 topics × 3 runs, medians, Claude Sonnet 5.5):

Ask for HTML directly Answer me with HTML
Output tokens 5,341 870 6.1× fewer
Time 33 s 12 s 2.8× faster
Cost per answer $0.092 $0.067 27% cheaper

The same TCP question answered both ways

One run from an earlier benchmark in a heavily loaded setup: same prompt, same model, and both pages are usable. This run took 9,351 output tokens for the plain page and 899 with the skill.

The speed-up holds in every setup. The cost depends on how much context your setup loads: the skill adds two short turns, and every turn re-reads the context. In a heavy setup (about 51,000 tokens of tools, rules and skills) the two extra turns cost more than the saved tokens, and we measured the skill about 20% more expensive.

Explainer videos show a bigger gap. We asked for the TCP handshake as a 3Blue1Brown-style video, no voice, both ways:

Write the video page by hand am video
Output tokens 27,839 1,566 17.8× fewer
Time 202 s 17 s 11.8× faster

That is one topic and five runs (3 by hand, 2 with am video, medians), so read it as a rough size, not a precise ratio. Per-topic numbers and the script to reproduce them are in bench/.

Install

You need Node.js 20 or newer. There is no npm install step. The CLI is bundled inside the skill.

Let your agent install it (recommended)

Paste this into Claude Code, Codex, Cursor, OpenCode or any other agent:

Install Answer me with HTML: read https://raw.githubusercontent.com/QingYunA/answer-me-with-html/main/INSTALL.md and follow it.

INSTALL.md is written for agents. It installs the plugin in Claude Code and the skill in other agents, keeps an existing install, asks you nothing, and ends with one report after checking the result with the TCP three-way handshake page. If your agent cannot open links, use npx -y skills add QingYunA/answer-me-with-html -g -y -a (for Claude Code, -a claude-code).

Claude Code plugin

Run this inside Claude Code:

/plugin marketplace add QingYunA/answer-me-with-html
/plugin install answer-me-with-html@answer-me-with-html

One command

npx skills add QingYunA/answer-me-with-html

It asks which agents to install into. The installer, vercel-labs/skills, supports more than 70 agents.

Manual install

Copy the skills/answer-me-with-html folder into your agent's skill folder. For Claude Code:

git clone --depth 1 https://github.com/QingYunA/answer-me-with-html.git /tmp/answer-me-with-html
cp -R /tmp/answer-me-with-html/skills/answer-me-with-html ~/.claude/skills/answer-me-with-html

Skill folders for other agents: Codex ~/.codex/skills/, Cursor ~/.cursor/skills/, OpenCode ~/.config/opencode/skill/.

No setup is needed after install.

What you ask, what you get

You ask You get
"Explain the TCP three-way handshake" A sequence diagram, a state diagram and a flag table
"How are the modules in this repo organized?" A folder tree plus a call graph
"Redis or Memcached?" A comparison table with ✓ and ✗, then a verdict
"What's wrong with this paragraph?" Each sentence annotated, with the problem words and fixes
"How did Kubernetes come about?" A timeline with the key moments highlighted
"How do I show hidden files with ls?" No page. A one-line question gets a one-line answer

The agent decides when a page is worth it: related concepts, multi-step flows, multi-way comparisons. You can also just say "explain it in HTML".

Pages are saved in ~/.answer-me-with-html/pages/. The buttons in the top-right corner switch the theme and light/dark mode, and copy the Markdown that produced the page.

Explainer videos (3Blue1Brown style)

Karpathy's ladder for understanding LLM output ends with explainer videos. Ask for one: "make a 3b1b-style video on the TCP handshake".

Four frames from a generated explainer video in the blueprint style: title card, a sequence diagram with the Server highlighted, a flow diagram, and a comparison table

The agent writes the same kind of draft as for a page, plus one line of narration per beat. Nothing else:

## Both sides wait
```sequence
Client -> Server: SYN
Server -> Client: SYN-ACK
```
> The client sends a SYN to ask for a connection.
> The [Server] answers with a SYN-ACK.

am video turns it into a player page:

  • Built step by step. When the Nth line of narration plays, the Nth step of the diagram appears. Arrows draw themselves. If there are more lines than steps, the extra lines at the start act as an intro.
  • Camera focus. [Server] in the narration pushes the camera toward that node and highlights it. The diagram never leaves the frame.
  • Objects carry over. A node with the same name in the next scene glides to its new place instead of cutting.
  • Narration. It uses ElevenLabs if ELEVENLABS_API_KEY is set (optional; setup guide), the system voice otherwise, and captions only if neither exists. Each beat lasts as long as its audio, so picture and voice stay in sync.
  • One file. The page has the audio inside and plays offline. Add --mp4 for a 1080p video file. This needs Chrome, ffmpeg and Node.js 22+ on your machine.

The draft for the example video (examples/video-tcp.en.md) is 1.3 KB, a few hundred output tokens. The agent only makes videos when you ask. A local voice server, themes and export time are in video details. Full syntax: am help video.

Settings

Change settings with a slash command. There are no config files to edit by hand.

Where How
Claude Code (plugin install) /answer-me-with-html:config asks what to change. /answer-me-with-html:config open off changes it directly
Any agent /answer-me-with-html config open off, or just say "stop opening the browser"
Terminal am config to view, am config set open off to change, am config reset to restore defaults
Key Default What it does
open on Open each page in the browser after it is made. Turn it off if pop-ups interrupt you
always on Always-on mode (see below). Only matters when the always-on plugin is installed
theme auto Default theme: auto (paper for long text, blueprint for diagrams), blueprint, shadcn, paper, or your own theme
mode auto Default color mode: auto, light or dark
style 80 Writing check: off, 80 (warn only) or strict (refuse to render)
update_check on Check GitHub for a new version once a week and mention it. Never updates by itself
voice auto Video narration: auto (ElevenLabs if ELEVENLABS_API_KEY is set, else system voice), elevenlabs, local, system or off

Settings live in ~/.answer-me-with-html/config.json. A theme written in a draft beats the default. --open and --no-open affect one run only.

Always-on mode (optional)

By default, the agent makes a page only for questions that need one. If you want a page with every conclusion, turn on always-on mode.

The agent then gets a short reminder each turn (about 90 tokens). Whenever it gives a conclusion, summary, plan or comparison, even a short one, it adds a small page with 2 to 4 panels and puts the path at the end of the reply. These pages never pop open, so they don't interrupt you. Casual chat and replies with no conclusion stay as they are. Claude Code makes no pages in plan mode.

Claude Code: install one more plugin.

/plugin marketplace add QingYunA/answer-me-with-html
/plugin install answer-me-with-html-always@answer-me-with-html

Pause it with /answer-me-with-html:config always off. You don't need to uninstall. That command exists only in the plugin install; if you installed the skill with npx skills, use /answer-me-with-html config always off.

Other agents: paste this to your agent so it writes the rule into its own rules file, such as AGENTS.md:

Turn on always-on mode for Answer me with HTML: add a global rule — "[answer-me-with-html always-on] Whenever a reply gives a conclusion, summary, plan, comparison, review or explanation, even a short one, also make a page with the answer-me-with-html skill (2 to 4 panels for routine answers), render it with --no-open before you write the reply, and end the reply with the page path. Skip casual chat, one- or two-sentence replies with no conclusion, pure command output, and requests for plain text."

Updating and cleaning up

Updates are manual. Once a week a background check reads the latest version number from GitHub (nothing about you or your pages is sent), and the agent mentions a new version when there is one. Turn it off with /answer-me-with-html:config update_check off. To update, run npx skills update answer-me-with-html -y, or tell your agent "update answer-me-with-html".

Pages, videos and the narration cache build up in ~/.answer-me-with-html/. When the folder gets large, the agent asks before cleaning, and nothing is deleted without your OK. Say "clean up the pages", or run am clean --dry-run to preview.

Update steps for other install methods and every am clean option: reference.

Background

Andrej Karpathy posted that as LLMs do more of the work, keeping up with their output becomes the hard part. A diagram or a web page is far easier to take in than a long block of text.

I tried asking agents to answer in HTML directly. The pages were good, but slow: a decent page took a minute or two, mostly hundreds of lines of CSS that were nearly the same every time. Diagrams were worse. The model had to compute SVG coordinates by hand, and arrows often pointed at nothing.

So Answer me with HTML takes that work away from the model. The model writes content. The CLI handles layout, color and drawing.

How it works

This is all the model writes:

---
title: TCP three-way handshake
---
## A Three-way handshake {span=2}
```sequence num
Client -> Server: SYN, seq=x
Server -> Client: SYN+ACK, seq=y, ack=x+1
Client -> Server: ACK, ack=y+1
note Client, Server: ESTABLISHED
```

## C State changes {span=2}
```flow LR
(CLOSED) -> LISTEN: passive open
LISTEN -> SYN_RCVD: get SYN / send SYN+ACK
SYN_RCVD -> *ESTABLISHED: get ACK
```

The CLI does the rest. It picks the template, places the panels, applies the theme, lays out the flow chart with dagre and spaces the sequence diagram by label width. The full draft, examples/tcp.en.md, becomes this page:

The TCP example page

Features

  • Layout by code: Panel placement and diagram coordinates are computed, not guessed. Labels don't get cut off, and there are no gaps in the grid.
  • Fixes its own mistakes: When a draft has an error, the CLI returns the line number, the component and a correct example. The agent fixes it in one try.
  • Three themes: blueprint looks like an engineering drawing, shadcn uses clean cards, and paper is set for long reading. By default the CLI picks paper for long text and blueprint for diagrams. All have light and dark modes, and you can add your own.
  • One file, no dependencies: Each page is a single .html with no CDN links or web fonts. It opens offline and is easy to share.
  • Writing check: Drafts are checked against rules adapted from ASD-STE100: long sentences, wordy phrases, passive voice. It only warns unless you ask for strict mode.
  • Keeps its source: Every page embeds the Markdown that made it. Click "Copy source" to get it back.

Blueprint theme

shadcn theme, dark

Blueprint theme (examples/ste100.md)

shadcn theme, dark mode

Components

The agent picks a component by the shape of the information:

Component Good for
flow Architecture, call chains, decision branches. Auto layout, with groups, decisions and databases
sequence Messages going back and forth between several parties over time
tree Folders, modules, taxonomies
timeline History, releases, phases
limits A value against its limit
annot Word-by-word notes on a sentence
kv Metadata, a drawing's title block
callout A conclusion, a tip, a warning
Table Multi-way comparison. Write ok / no / warn in a cell to get ✓ ✗ !

The draft format (frontmatter, span, rows, raw html / svg blocks) and how to call the CLI without an agent are in the reference. Full syntax for a component: am help .

The STE writing check

ASD-STE100 is a controlled form of English first used for aircraft maintenance manuals. Its rules are concrete: keep sentences short, give each word one meaning, write steps as commands. Karpathy noted that asking an LLM to follow these rules makes its writing much easier to read.

Answer me with HTML turns the parts a machine can check into an English and Chinese rule set, and runs it on every render:

  • Length: Steps stay under 20 English words or 35 Chinese characters. Descriptions stay under 25 words or 45 characters. Paragraphs have at most 6 sentences.
  • Words: Prefer common words: "use", not "utilize"; "before", not "prior to". In Chinese, drop empty verbs: write 优化, not 进行优化.
  • Chinese vocabulary: Common typos (登陆 → 登录), vague quantities (尽快, 若干, 大概, 多次), 以上 / 以下 / 以内 after a number (the endpoint is ambiguous), and one-meaning-one-word choices (单击 → 点击, 键入 → 输入, 入参 → 参数). Only the entries that are almost never wrong, taken from Simplified Technical Chinese, a controlled Chinese modelled on the STE method.
  • Style: Flags English passive voice, three or more 的 in one sentence, and stock phrases such as 赋能 and 闭环.
  • Japanese: A draft with kana is treated as Japanese: the page buttons are in Japanese and the page gets lang="ja". Only the length rules apply, with the Chinese character limits. Write lang: ja in the draft to force it.

Set the strictness with /answer-me-with-html:config style strict, or per page with style: in the draft.

Development

git clone https://github.com/QingYunA/answer-me-with-html.git && cd answer-me-with-html
npm install
npm test          # run the tests
AM_E2E=1 npm test # also run end-to-end video tests (system TTS, Chrome, ffmpeg)
npm run smoke:install # install for real with npx skills and validate the plugin manifests (needs network)
npm run build     # after changing src/, rebuild skills/answer-me-with-html/scripts/am.mjs
npm run snapshot  # compare rendered HTML with origin/main (refactors must not change it)

Maintainer conventions (generated bundle, page format, snapshot checks, reviewing PRs, releasing) are in CONTRIBUTING.md.

There are two runtime dependencies: marked parses Markdown and @dagrejs/dagre lays out flow charts. Both are bundled into am.mjs.

Community

Discussion and feedback also happen on LINUX DO, a Chinese-language developer forum.

Star History

Star History Chart

License

MIT