Skip to content
Ghostchars

Blog

Ghostchars is on npm: a CLI, a CI gate, and an MCP server

Four commands that share their switches with this website, a gate that holds a whole tree to a bar in CI, and a skill and an MCP server so an agent runs it unasked. Nothing here calls home.

By Kris Vinters · Published · Updated · 5 minute read

What ships

The Ghostchars command-line tool is on npm. It was published on 2026-09-06 as ghostchars@1.0.0, and it installs from its own page on this site, which carries the registry link only while the release is actually there. It needs Node 22 or newer. The package has zero runtime dependencies, and the whole thing is a launcher and three bundles: the CLI, a document reader, and an MCP server. Only the first one loads unless you name a document or start the server.

Four commands, and they share their switches with this website and the Mac app. clean writes the cleaned text to stdout, or rewrites a file in place with -i, atomically. inspect lists every suspicious character with its line, column, codepoint and Unicode name, and changes nothing. report measures the habits that make a draft read as machine written and is not a detector: no probability, no score, no verdict. check holds a whole tree to a bar, which is the next section.

npx ghostchars clean -i notes.md          # clean a file in place, atomically
cat draft.md | npx ghostchars inspect     # what is hiding: line:col, codepoint, name
npx ghostchars check --bar text           # the gate: exit 1 if the tree is dirty
npx ghostchars report draft.md            # habits that make a draft read as machine written
npx ghostchars clean -i essay.docx        # documents too: .docx, .odt, .html

The switches themselves are grouped into three bars, so a person or an agent names one word instead of twelve flags. default strips only what is invisible. text adds spaces and typographic punctuation on top, -s --punct. And max turns on all eleven switches, including confusables and Unicode normalisation, and it is lossy: do not run it over text you did not write yourself.

clean and inspect also open .docx, .odt and .html and handle them structurally, text nodes only, rather than as a stream of bytes. Hidden runs, comments and tracked changes are reported and never removed, because deleting them changes the document and that stays a human call. PDF, images, spreadsheets and Google Docs are out of scope.

The gate, and the SARIF upload

check is the command for CI and for a pre-commit hook. With no paths named it finds files with git ls-files, so it honours the ignore rules the repository already has, and it prints every file that would change with exact character positions. A ghostchars.json at the root names which files are held to which bar, and pins a file that changes by exactly N characters on purpose, enforced in both directions: a file that was quietly fixed after being pinned fails the run instead of passing silently.

npx -y ghostchars check --bar text

Never pass --fix in CI. It cleans the failing files in place and re-checks, which turns a red build green without anyone reading what changed.

With --sarif, the same run prints as a SARIF 2.1.0 log instead, one result per offending character, and GitHub code scanning shows the findings inline on the pull request.

- run: npx -y ghostchars check --bar text --sarif > ghostchars.sarif || true
- uses: github/codeql-action/upload-sarif@v3
  with:
    sarif_file: ghostchars.sarif

The skill, the hook, and the MCP server

The package ships the files that let an agent run the gate without being asked each time, and none of them is an installer. The skill is one directory, holding SKILL.md and two reference files, and the same copy is read by Claude Code, Codex, Cursor, VS Code with Copilot and Gemini CLI. It tells the agent when to run the tool and, more usefully, what to do with each kind of finding, including the one a flattened em dash leaves behind.

mcp runs a Model Context Protocol server on stdin and stdout, so an agent calls clean_text, inspect_text, style_report and check_paths as typed tools instead of shelling out and parsing the reply back into JSON. All four read only. A fifth tool, clean_files, writes to disk, and it stays unregistered rather than refusing when the server starts without --allow-write, because an agent should never discover a tool that will always fail.

The Claude Code hook runs after every file the agent writes or edits, holds that one file to the text bar, and hands a finding back as feedback rather than rewriting anything itself: the model gets the chance to remove the character or rewrite the sentence in the same turn. The pre-commit block does the equivalent job at the commit, over the staged text files, so nothing dirty gets committed by anyone, and like the CI step it never passes --fix either.

Privacy

There is no network code in the package at all: no telemetry, no update check, no licence check, no account. Nothing you clean, inspect or report on leaves the machine it runs on, whether that machine is your own terminal, a CI runner, or the sandbox an agent is calling it from. The only bytes that leave the process are the ones it prints, and the files you asked it to rewrite.

You do not have to take that on trust. The engine, the CLI and the parity tooling that holds the two to the same answers are open source under the MIT licence, at github.com/ghostchars/ghostchars.

The same engine, elsewhere

The same engine runs in two more places, for whenever a terminal is not where you write. The Mac app puts it behind one keystroke, in every app you type in. The browser extension puts it in a popup and a right-click menu, free and unlimited, in Chrome and the browsers built on it.

Read the full command reference, and install it · Download the Mac app · Get the browser extension

Corrections

If something here is wrong, write to privacy@ghostchars.com and it will be corrected in place, with the update date at the top of the post moved to match.

All posts

Anonymous usage stats?

Ghostchars processes your text in the browser and never uploads it. We would like to count page views with a self-hosted, cookie-light analytics endpoint. No content, ever.