Translate the web and your PDFs with a language model running on your own machine. No cloud, no API key, no telemetry.
Select text on any page — or in a PDF — and the translation streams in place. Inference happens on a local Ollama server, so nothing you read ever leaves your computer.
Real time, unedited, against a local qwen3.
- You read things you cannot paste into a cloud translator — unpublished drafts, contracts, medical records, internal documents, anything under NDA. Cloud translators are excellent and completely unusable for this.
- You want the whole page, not one sentence at a time. Bilingual translation puts each paragraph's translation directly under the original, so a sentence that looks wrong can be checked against the source instead of taken on faith.
- You read PDFs, not just web pages. Papers, specs, whitepapers. OpenRead
opens them in a bundled PDF.js viewer where the same selection translator
works on the rendered text — including
file://PDFs already on your disk. - You want Traditional Chinese that reads like Taiwan. Most tools translate
into Simplified and convert character-by-character, which produces
界面,公裡, and數據庫where a Taiwanese reader expects介面,公里, and資料庫. OpenRead converts at the phrase level with OpenCCs2twp, and the conversion is tested against chunk-boundary corruption. - You want to keep what you read. One tap saves any translated selection into your Obsidian vault as Markdown.
Download the latest release, unzip it, then open chrome://extensions,
enable Developer mode, click Load unpacked, and select the unzipped
folder.
Or build from source
pnpm install
pnpm buildThen load .output/chrome-mv3 as an unpacked extension.
OpenRead needs a local model server. This is the only setup step, and it is a one-time cost:
-
Pull the default model:
ollama pull qwen3(chosen by benchmark, not by taste). -
Let the extension talk to it. Ollama only answers requests from origins it knows, and a browser extension is not one of them by default. Set
OLLAMA_ORIGINS=chrome-extension://*and restart Ollama:OS Command macOS launchctl setenv OLLAMA_ORIGINS "chrome-extension://*", then restart OllamaLinux add it to the systemd unit or your shell env, then restart Ollama Windows set a user environment variable OLLAMA_ORIGINS=chrome-extension://*, restartSkip this and every translation fails with a
403. The extension says so in as many words, and links back here.
Open the toolbar popup to confirm the connection, pick a model, and choose a target language.
Whole-page bilingual translation, live qwen3. The navigation, table of
contents and account links are deliberately untouched — only the article is
translated.
- Whole page — click Translate this page in the popup, or press
Ctrl+Shift+G. Each paragraph gets its translation underneath it, the page fills from the top, and a badge in the corner counts progress and offers a Stop. Press the shortcut again to remove every translation and get the original page back. - One selection — select text, click the floating 文, and the translation streams into a panel.
- PDFs — open any
.pdf; OpenRead redirects it into the bundled PDF.js viewer, where selection works exactly the same. - Keyboard — select with Shift+Arrow or Ctrl+A and press
Ctrl+Shift+Y(remappable atchrome://extensions/shortcuts). Escape closes the panel.
| Developer docs (MDN) | Research PDF (PDF.js viewer) |
|---|---|
![]() |
![]() |
Every screenshot on this page is a real end-to-end run against a local
qwen3 — the built extension loaded into Chrome, not a mock-up.
Reading is half the loop. Once a translation streams in, a + Save to
Obsidian button drops a Markdown note — original, translation, and a
machine-readable YAML header — into your vault through an obsidian://new URI.
No extra permissions, no server; notes too large for a protocol-handler URL fall
back to the clipboard.
| Translate + one-tap capture on any page | …lands as a note in your Obsidian vault |
|---|---|
![]() |
![]() |
Every note is written status: raw. That header is a deliberate handoff
contract: a stronger downstream model can query the unprocessed captures,
synthesize them, and flip the flag. OpenRead does the cheap, reliable part
on-device and defers the expensive part, rather than re-implementing a knowledge
base it has no business owning.
Set your vault, capture folder, and the optional enrichment toggle in the popup; leave the vault blank to use whichever vault is currently open.
The selected text is sent to one place: the Ollama server at the URL you
configured, on your own machine. There is no account, no API key, no analytics,
and no remote endpoint anywhere in the code. The server URL lives in
chrome.storage, is read only by the background worker, and never travels over
the extension's message bus. Permissions are storage and activeTab —
v1 declared two more it never used.
Everything above is the product. The rest of this file is the engineering, which is the part this repository actually exists to show: making streaming LLM output reliable, and proving it — with translation as the vehicle.
An LLM told to "translate this" will happily also emit a preamble
(Sure, here is the translation:), think out loud (The user wants…), echo the
input back, wrap the output in quotes, or — for a Traditional-Chinese target —
leak Simplified characters. In a streaming UI these artifacts land on screen
before you can react.
OpenRead treats that as an engineering problem with a measurable target. The cleanup logic is a pure, dependency-free core, unit-tested in isolation and scored by an offline eval harness so improvements are quantified, not vibes.
pnpm eval replays a curated set of real failure modes through the shipped
streaming pipeline and reports before/after rates. Fully offline and
deterministic — no Ollama server, no network — so the numbers are reproducible
in CI. Deltas are replayed in 3-character slices, so chunk boundaries land
mid-artifact exactly as they do in a live stream.
| Metric | Before | After | Reduction |
|---|---|---|---|
| Preamble / thinking leakage | 34.8% | 0.0% | 100% |
| Input echo | 17.4% | 0.0% | 100% |
| Simplified-character leakage (Traditional targets) | 42.9% | 0.0% | 100% |
Measured over 23 curated fixtures (21 Traditional-Chinese targets). Regenerate
with pnpm eval; full report in eval/RESULTS.md.
309 unit tests cover everything with real behaviour (pnpm test:cov): the
pure core at 100% function / 97% line coverage, the selection and capture UI
driven through jsdom with a stubbed extension port, and the background worker —
which owns cancellation, error translation and PDF routing. Overall: 92%
function, 95% line.
The offline eval freezes model output to score the pipeline; pnpm bench
asks the opposite question against live models: which model should you run,
and what does each design choice cost? 27 curated EN→zh-TW fixtures with
Taiwan-convention references × 4 models × 2 prompt conditions, streamed
through the exact shipped pipeline and scored with sacrebleu-cross-validated
chrF, artifact detectors, latency probes, and a schema-constrained LLM judge —
itself calibrated against 40 blind human labels: quadratic-weighted Cohen's κ
0.53 on adequacy (moderate — usable), 0.21–0.27 on fluency/localization (weak
— so quality claims lean on chrF + adequacy, and the judge's localization
scores are treated as an upper bound; eval/AGREEMENT.md).
| Model | chrF ↑ | TTFT-UI p50 | Tokens/s | Verdict |
|---|---|---|---|---|
| qwen3 (default) | 46.4 | 451 ms | 48 | best quality/latency balance |
| qwen3.5 | 43.4 | 730 ms | 42 | no chrF edge, 1.6× the wait |
| llama3.1 | 31.9 | 532 ms | 49 | fast, but ~13 chrF behind |
| deepseek-r1:8b | 36.4 | 6,353 ms | — | 6-second "thinking tax" — wrong workload |
Engineered-prompt condition, seed 42; full tables in
eval/BENCHMARK-RESULTS.md, methodology and
limitations in docs/BENCHMARK.md.
Three findings worth calling out:
- The benchmark caught a product-breaking bug. Through Ollama's
OpenAI-compat endpoint, reasoning models can spend the entire generation
on hidden chain-of-thought — one fixture: 99 s, 4,055 tokens, zero visible
characters. The client now uses the native
/api/chatwiththink: false(same fixture: 1.6 s). - The reliability layer stopped being a tradeoff once it was fixed. It
zeroes preamble on dirty outputs and deepseek-r1's Simplified leakage
(14.8% → 0%), and now adds chrF on six of eight model × prompt cells —
llama3.1's naive prompt went from −0.3 to +1.0 after the streaming
assembler stopped flushing mid-artifact. Its remaining price is ~200 ms of
first paint. Same recorded generations, re-scored through the current
pipeline (
pnpm bench -- --repipe), so the change is the code, not sampling. - The eval was scoring a function the product does not call.
pnpm evalrancleanTranslationOutputand called it "the exact transform the production pipeline applies" — but the product streams, throughStreamAssembler, which never touched it. Replayed through the real path, the reported 0% preamble and 0% echo were actually 8.7% and 8.7%. Both leaks traced to one cause: the reluctant buffer's 12 characters is shorter than the artifacts it exists to catch, soHere is the translation: …flushed atHere is theand the rest streamed to the panel. The buffer is now adaptive and every harness replays the shipped assembler. A convenient function is how an eval starts lying about the thing it is meant to prove.
- Reliability layer (
src/core/sanitize.ts) — anchored preamble/thinking filters, echo removal, quote unwrapping. - Streaming assembler (
src/core/stream.ts) — a "reluctant buffer" holds only the opening tokens (where preamble hides) so the translation still paints fast, then streams the rest straight through. The hold is adaptive: it extends only while something is actually resolving — a preamble that has not reached its colon, an unclosed<think>block, an echo of the selection still arriving — and is capped so first paint cannot stall. Measured on real qwen3 output, clean translations are held for exactly as many characters as before; only artifact-shaped openings wait longer. - Taiwan localization (
src/core/zh-convert.ts) — OpenCCs2twpphrase-level Simplified→Traditional conversion, replacing v1's hand-rolled character map that corrupted界面→界麵and公里→公裡. Because it maps phrases, converting each stream chunk on its own mistranslates any phrase a chunk boundary splits (数据+库→數據庫, not資料庫), so the transform holds the ambiguous tail back until enough context arrives — the streamed result is byte-identical to converting the finished text in one call. - Verified in a browser, not only in jsdom
(
e2e/fullpage.mjs) — every unit test here runs in jsdom, which cannot tell you whether an extension loads, whether a content script is injected, or whether the service worker is awake when a message arrives. Three shipped defects (2.2.11–2.2.13) were invisible until the built extension was loaded into Chrome.pnpm e2e:pagedoes that loading over CDP, drives the real popup message against a live page and a real local model, and asserts the properties that matter: translations land, the original survives, chrome under the length floor is skipped, and toggling again restores the page byte for byte. It is deliberately not in CI — GitHub's runners have no GPU and no model, and a translation harness that stubs the model is measuring the stub. - Whole-page translation as a queue, not a flood
(
src/ui/fullpage.ts+src/ui/blocks.ts) — Ollama serves one generation per model, so firing fifty parallel requests only builds a queue in arrival order, which is not the order anyone reads in. Two requests stay in flight and the page fills top-down. Block selection takes the leaf-most prose element (aliwrapping apis translated once, not twice), scopes to the page'smain/articlelandmark when it declares one, drops navigational chrome, refuses to descend intocode/pre, honourstranslate="no"and.notranslate, and skips anything already in the target language through the sameshouldBypassAIshort-circuit selection uses. Measured on real pages: Wikipedia's article for Ollama goes from 325 candidate blocks to 48, and the first block translated changes from "Current events" to the opening sentence. Translations are appended, never substituted: a local 8B model is good, not perfect, and a reader has to be able to check a sentence that looks wrong. - Keyboard-first, not keyboard-afterthought — text selected with
Shift+Arrow or Ctrl+A offers the 文 button just as a mouse selection does, the
button takes Enter/Space, Escape dismisses the panel, and
Ctrl+Shift+Ytranslates the selection without touching the icon at all. The panel is a namedrole="dialog"whose content is anaria-liveregion, so a streamed translation is actually announced. - Same-language short-circuit
(
src/core/language.ts) — script detection skips the API entirely when a selection is already in the target language (zero latency, zero cost). The Simplified/Traditional marker sets are derived from the OpenCC dictionaries (pnpm gen:markers), not hand-written: the hand-written lists both flagged shared characters like 系 and 游 as Simplified and missed common ones like 发 and 时. - Cancellation-safe streaming
(
src/api/ollama.ts+src/entrypoints/background.ts) — each request owns anAbortController; a new selection or a closed panel aborts the in-flight stream with no shared mutable state to race on. - Reasoning-model safe — the client uses Ollama's native
/api/chatwiththink: falsebecause the benchmark caught the OpenAI-compat endpoint burning entire generations as hidden reasoning with zero visible output on qwen3-family and deepseek-r1 models (docs/BENCHMARK.md§6). Requires Ollama ≥ 0.9.
See docs/ARCHITECTURE.md for the module map and the
streaming sequence diagram.
Optionally, a small local model can pre-label a capture with a title, summary, and tags. That pipeline is defense-in-depth, and every layer is measured:
| Layer | Evidence |
|---|---|
Schema-constrained decoding (Ollama format) |
took the one imperfect model (deepseek-r1) from 93.3% → 100% usable metadata at zero latency cost — live study over 4 models × 16 excerpts (eval/STRUCTURED-RESULTS.md) |
Tolerant parseEnrichResponse |
salvages 71.4% vs naive parsing's 42.9% on an archive of 14 hostile reply shapes from older/thinking models (eval/CAPTURE-RESULTS.md); today it mostly does content hygiene — length caps, tag normalisation |
status: raw handoff |
even a perfect-looking label is garnish; the raw capture stays the source of truth |
The live study also produced an honest negative result: with the shipped prompt, temperature 0, and thinking disabled, modern small models emit clean JSON ~100% of the time — the dramatic salvage rates belong to older model generations. Enrichment stays off by default anyway: it adds a model round-trip per capture, and reasoning-class models take ~45 s to label a paragraph (measured), which no capture UX survives.
v1 was ~1,500 lines of untyped JavaScript. Each of these was found by rebuilding it, and each is why a corresponding test exists:
viewer_init.jsloadedpdf.worker.js; the file ispdf.worker.mjs. The PDF worker never started.utils/zh-map.jsadvertised "~2,800 pairs"; the string was ~200 characters repeated four or five times.- Unconditional
面→麵,里→裡,台→臺substitution corrupted界面→界麵and公里→公裡. OpenCC replaced it. - The manifest declared
scriptinganddeclarativeNetRequestand used neither — a store-review red flag, and permissions a privacy-first extension has no business asking for. content.jsandpdf-integration.jswere ~90% copy-paste of each other. They are now one module.
pnpm dev # HMR dev build (Chrome); pnpm dev:firefox for Firefox
pnpm test # Vitest unit suite
pnpm test:cov # …with coverage
pnpm eval # reliability eval -> eval/RESULTS.md
pnpm eval:capture # capture-enrichment eval -> eval/CAPTURE-RESULTS.md
pnpm bench # live model benchmark (needs Ollama) -> eval/BENCHMARK-RESULTS.md
pnpm e2e:page # whole-page translation in a real Chrome (needs Ollama)
pnpm compile # tsc --noEmit (strict)
pnpm lint # ESLint
pnpm build # production build -> .output/chrome-mv3Releases are cut by pushing a tag: CI re-runs every gate, extracts the notes
from CHANGELOG.md, and attaches the built zip.
See CONTRIBUTING.md for the full workflow.
TypeScript (strict) · WXT · Vitest · ESLint + Prettier · OpenCC · Ollama · GitHub Actions






