- Home
- Plugins
- Tools & Capabilities
- dsh-doc-router
dsh-doc-router
dantwz7/dsh-doc-router
DeepSeek Harness document-routing tools: sniff a file's real format and a PDF's page layout, then extract with the right pipeline — layout-aware Markdown for multi-column papers, MarkItDown for everything else, images for scans.
Install
dsh plugin --profile web add github:dantwz7/dsh-doc-routerREADME
dsh-doc-router
Route before you read. A DeepSeek Harness plugin that inspects a document's real format and a PDF's page layout, then hands it to the extraction pipeline it deserves — converting multi-column papers with layout awareness instead of shredding them.
Why · What you get · Install · Quick start · How the routing works · Configuration · Evidence · Troubleshooting · Development · Roadmap · License
Why this exists
Default converters get multi-column papers wrong, and the failure is not subtle.
On one Nature Letter (nmat4749.pdf, two columns, 7 pages), measured with
PyMuPDF 1.28.2 / MarkItDown 0.2.0:
| Pipeline | Output | Multi-column handling |
|---|---|---|
pymupdf4llm (layout-aware) | 6 533 chars for page 1 | ✅ columns merged in reading order; headings → #, superscripts → <sup>, _in situ_ italics kept |
pypdf | 30 632 chars, whole document | ⚠️ text is readable but structure is gone |
| MarkItDown | 45 266 chars, whole document | ❌ body text fragmented into hundreds of | … | table rows |
The same failure is reproducible from a clone. On this repository's
two-column-reading-order.pdf
and three-column.pdf
— synthetic pages whose columns are individually marked — MarkItDown 0.1.5 reads
straight down every column at once:
| Fixture | pdf_markdown | MarkItDown 0.1.5 |
|---|---|---|
two-column-reading-order.pdf | L01…L15, then R01…R15 | L01 R01 L02 R02 … — 29 column alternations |
three-column.pdf | L01…L15, M01…M15, then R01…R15 | L01 M01 R01 L02 … — 44 column alternations |
That is the whole argument in one line: one pipeline keeps the columns, the other
interleaves them. npm test asserts the reading order in both pdf_markdown cells.
Asking a model to "just pick a converter" therefore produces quietly mangled source material, and the damage looks like success. The fix is not a better single converter — it is routing: look at the file first, then choose the pipeline.
What you get
| Artifact | Purpose |
|---|---|
doc_route tool | Classifies a file and names the pipeline to use: format, page count, whether a text layer exists, single- vs multi-column, plus the per-page evidence behind the verdict |
pdf_markdown tool | Layout-aware PDF → Markdown via PyMuPDF; supports a page range and writing to disk |
doc-routing skill | Registered into the runtime skill layer, so it applies to every workspace without a per-workspace skills/ directory. Tells the model to route before reading |
Every verdict is deterministic — no model call, no network, no token cost.
Install
Requirements
| DSH | >=0.2.0-rc.1 <0.3.0-0 |
| Node | `^22.19.0 |
| Python | 3.9+ — needed only for PDF layout detection and pdf_markdown |
| PyMuPDF | pip install pymupdf4llm — optional; see license |
The interpreter is looked up in this order, first hit wins:
config.pythonPath- the
DOC_ROUTER_PYTHONenvironment variable - the runtime DSH ships
- a runtime under
~/.dsh/dsh-runtimes pythononPATH
An explicit choice is taken verbatim, so a bare command name like python3 is
allowed — and a wrong path is reported as the path you named, not as a mystery.
From npm (recommended)
plugin_manager action=install_bundle target="dsh-doc-router@0.6.0"
Pin the exact version. DSH's package manager applies a minimum-release-age policy, and installing by bare name can silently resolve to an older build. Restart DSH afterwards.
From the DSH CLI
dsh plugin --profile <YOUR-PROFILE> add dsh-doc-router@0.6.0
<YOUR-PROFILE> must be the profile DSH actually boots. plugin_manager above
targets the active profile for you, so prefer it if you are unsure. Two things this
subcommand does not do:
- A wrong
--profilename is not an error — it creates one. The profile is initialized as a new, empty profile and the package is installed there, so the plugin never reaches the profile you are running. The active profile is the directory under$DSH_HOME/profileswhosepackage.jsonlists your plugins. --helpis not a dry run.dsh plugin --profile <name> --helpinitializes that profile before printing pnpm's help.
From a checkout instead
plugin_manager action=install_bundle target="file:<ABSOLUTE-PATH-TO-REPO>"
A file: install is a copy, so editing the source does not take effect: bump
the version, remove the bundle, reinstall and restart DSH. See
CONTRIBUTING.md.
Enable layout detection
<the python the plugin resolved> -m pip install pymupdf4llm
Without PyMuPDF the plugin still loads and doc_route still classifies by
format; only layout detection and pdf_markdown are disabled, and both say so
rather than failing obscurely.
Quick start
Route first, then convert. The examples below use this repository's own synthetic fixtures, so every line is reproducible from a clone.
1. Route it
doc_route({ path: "tests/fixtures/two-column.pdf" })
→ Routed tests/fixtures/two-column.pdf — format: pdf, pages: 1, text layer: yes, columns: multi-column (est 2)
Recommended pipeline: pdf_markdown
Multi-column body text (est. 2 columns) — markitdown shreds it into table fragments and pypdf drops the structure
If formulas, super/subscripts or digits must be exact, render that page to an image and check it
Probe detail: {"1":"0.91/2col"}
dsh-doc-router v0.6.0
The two advisory lines above are English because that is the default as of
0.4.0— see Configuration fornoteLanguage: 'zh'. The verdict line,recommend, andprobefields are language-neutral either way.
2. Convert it
pdf_markdown({ path: "tests/fixtures/two-column.pdf" })
→ Converted tests/fixtures/two-column.pdf (all 1 pages) — 3391 characters of Markdown. …
pdf_markdown({ path: "tests/fixtures/two-column.pdf", pages: "1", output: "paper.md" })
→ Converted tests/fixtures/two-column.pdf (pages 1) — wrote 3391 characters to paper.md.
Both tools accept workspace-relative paths. When inline output would exceed
maxChars, pdf_markdown truncates and says so instead of flooding the
context — pass output to get the whole document on disk.
Real verdicts, from this repository's fixtures
Every row below was produced by running doc_route on the checked-in fixture:
| Fixture | Verdict | Pipeline |
|---|---|---|
two-column.pdf | pdf · 1 page · text layer · multi-column (est 2) | pdf_markdown |
three-column.pdf | pdf · 1 page · text layer · multi-column (est 3) | pdf_markdown |
two-column-watermark.pdf | pdf · 1 page · text layer · multi-column (est 2 — not 3) | pdf_markdown |
unknown-columns.pdf | pdf · 3 pages · text layer · unknown | pdf_markdown (safe side) |
dense-script-column.pdf | pdf · 1 page · text layer · multi-column (est 2) | pdf_markdown |
single-column.pdf | pdf · 1 page · text layer · single-column | markitdown → pypdf |
scanned.pdf | pdf · 1 page · no text layer | render_to_png → read_image |
sample.docx | docx | markitdown |
sample.png | png | read_image |
sample.txt | text | read |
pdf_markdown refuses non-PDFs and names the pipeline to use instead:
pdf_markdown({ path: "tests/fixtures/sample.docx" })
→ pdf_markdown only converts PDFs — detected docx. Recommended pipeline: markitdown
(run doc_route for the full verdict)
[!NOTE] Advisory language. The verdict line,
recommend, andprobefields are language-neutral. Every human- or model-facing string followsnoteLanguage— the free-text advisory lines after the verdict,pdf_markdown's refusal message, and thedoc-routingskill body.enby default since0.4.0,zhfor the previous behaviour.
How the routing works
All decisions are deterministic: zero model calls, zero network, zero tokens.
| Decision | Basis |
|---|---|
| Format | The file header's magic bytes, never the extension. ZIP containers are opened to tell docx / xlsx / pptx / epub / odt apart; BOM'd and BOM-less UTF-16/UTF-32 plain text is decoded, with a control-character guard so a binary is not mistaken for text |
| Text layer | Fewer than 120 characters per page ⇒ treated as a scan |
| Single vs multi-column | Text intervals are merged per horizontal band and the segments separated by internal whitespace are counted. Any body page that is multi-column makes the document multi-column. Edge artefacts — a publisher's rotated margin stamp, headers, footers — are excluded: they are narrow but run the full page height, so they land in every band |
That last rule is a deliberate asymmetric bet: pymupdf4llm works fine on
single-column pages too, whereas a missed multi-column page is handed to a
converter that destroys the body text. One mistake is cheap; the other is not.
The same bet decides what happens when the column count cannot be measured —
a page whose body is one large block, or many fragments too short to be body text,
yields no usable band evidence. That is reported honestly as unknown and routed to
pdf_markdown, and is never described as single-column: guessing "single
column" is the expensive direction, so a file we cannot measure takes the cheap one.
"Too short to be body text" is measured in content, not codepoints. A block is
discarded only below 25 latin-equivalent codepoints, where an ideograph counts as
two — because a Chinese word is 1–2 ideographs while an English word is ~5 letters.
Counting raw codepoints instead made the filter several times more aggressive on
Chinese PDFs, and could throw away an entire column of real body text; see
dense-script-column.pdf.
The routes
| Verdict | Pipeline |
|---|---|
| multi-column + text layer | pdf_markdown |
| single-column + text layer | markitdown, falling back to pypdf |
column count unmeasurable (unknown) + text layer | pdf_markdown — the safe side |
| no text layer | render to PNG → read_image |
| docx / xlsx / pptx / epub / odt / csv / html / json / xml | markitdown |
| images | read_image |
| text / rtf | read — no conversion at all |
Why multi-modal reading is not on the normal path
Format and column count are geometry problems. Code is fast, exact and free; a single 150 dpi page image is roughly 0.5 MB and a thousand tokens, and must be rendered page by page. Images are therefore reserved for the three things code cannot do:
- No text layer — a scan.
- Exactly verifying formulas, superscripts, Greek letters and units — text
extraction gets these wrong in ways that matter (an image shows
Sr₃Al₂O₆correctly; text extraction can only giveSr3Al2O6). - Understanding figures, layout, or which caption belongs to which figure.
Three routing bugs worth not reintroducing
- "Count the text blocks on each side" → on a Nature first page the left column held only 2 blocks, so it was misread as single-column. Fixed by finding gutters per line instead.
- "Only look for the gutter down the middle of the page" → a 3-column Science layout has body text in the middle, so it was missed. Fixed by counting columns.
- "Take the majority across pages" → a single-column reference page and a figure page tied, tipping the verdict to single-column. Fixed to any body page multi-column ⇒ multi-column, ignoring pages under 300 characters.
Two further defects were found by a real-corpus trial and fixed in 0.2.1; both are
regression-tested and written up in CHANGELOG.md:
- A rotated full-height margin stamp was counted as an extra column, so a two-column paper reported three columns — 18 of 74 corpus files.
columns: unknownshared a branch withsingle-columnand was routed tomarkitdownwhile the note asserted "single-column". This was the one place where the plugin contradicted its own asymmetric bet.
Configuration
Every field is optional.
| Field | Default | Effect |
|---|---|---|
pythonPath | (unset) | Interpreter to spawn, taken verbatim — highest priority |
timeoutMs | 120000 | Probe timeout, milliseconds |
maxChars | 120000 | Inline Markdown cap; beyond it pdf_markdown truncates and says so |
noteLanguage | en | Language of every human- or model-facing string — advisory notes, the pdf_markdown refusal, and the skill body: en or zh |
This is the config block of the doc-router entry inserted by
cordis.patch.yml:
- insert:
- id: doc-router
name: dsh-doc-router
config:
pythonPath: C:\path\to\python.exe # optional; highest-priority interpreter
timeoutMs: 120000 # optional; probe timeout, ms
maxChars: 120000 # optional; inline Markdown cap
noteLanguage: en # optional; 'en' (default) or 'zh'
noteLanguage changes only human-facing text. The verdict, recommend, and
probe fields are language-neutral, and routing is identical in either language.
Evidence
The routing rules were exercised over a real corpus of 71 published PDFs plus one DOCX (ACS / Wiley / Science / Nature / Sci. Adv. / arXiv, 2–83 pages, Chinese directory and file names):
| Check | Result |
|---|---|
| Files routed | 72 |
| Multi-column detected | 47 — every one confirmed multi-column by eye; 0 missed |
| Single-column detected | 22 — none was actually multi-column |
| Column-count over-estimates | 18 → fixed in 0.2.1 |
unknown layouts | 2 → routed to the safe side in 0.2.1 |
| "Succeeded" but produced garbage | 0 |
| Verdicts changed by the short-block filter fix | 0 of 71 (whole corpus re-run, pre- vs post-fix) |
| Verdicts changed by the weighted-threshold fix | 0 of 71; 1 of 121 on a wider tree that includes 26 Chinese documents (a single-page figure, rendered and checked by eye) |
One of the 47 is a genuine three-column Science article (three equal 164pt
columns, abstract spanning the first two), so the est 3 branch is exercised by
real input and not only by the synthetic fixture.
On one page of a Science paper, the same input through both paths:
pdf_markdown | MarkItDown | |
|---|---|---|
| Characters | 5 959 | 11 987 |
| --- | table separators | 0 | 24 |
| Word spaces lost mid-word | 0 | 36 (e.g. constants.Theresultantcoherentgrowthof) |
| Cost and reliability, same corpus | Measured |
|---|---|
doc_route, per file | ~0.4 s |
pdf_markdown, typical paper | ~7.7 s |
pdf_markdown, 83-page paper | ~29.7 s |
| Path failures — 72 routes + 44 conversions on non-ASCII paths | 0 |
[!NOTE] The interleaving above is reproducible from a clone on a synthetic fixture. The table in this section uses real published papers because that is the harsher test — and those cannot be redistributed, so their exact numbers are recorded rather than reproducible. See tests/fixtures/README.md.
The corpus totals are, however, re-runnable in one command against a corpus you supply:
python -B scripts/validate-routing.py <corpus-root> --json out.jsonprints the per-file verdict table and the totals. The script ships with the repo (scripts/validate-routing.py) and imports the samelib/docprobe.pythe tool uses, so it cannot disagree with it. The recorded run is kept as a per-file manifest (path, size, pages, verdict, estimate, chars/page for all 71 files), so the numbers can be diffed rather than merely believed.
Troubleshooting
| What you see | What to do |
|---|---|
docprobe failed … Microsoft Store … exit code 9009 | The python on PATH is the Windows Store placeholder. Point pythonPath (or DOC_ROUTER_PYTHON) at a real interpreter. |
PyMuPDF is not installed | The interpreter the plugin resolved is not the one you installed PyMuPDF into. Set pythonPath explicitly to see which one it chose, or install into the packaged runtime under $DSH_HOME/dsh-runtimes/. Note that a DSH upgrade can overwrite that runtime's site-packages, so reinstall PyMuPDF afterwards. |
page selection '99' matched no pages; the document has 1 page(s) | The page range is validated against the real page count. Use "3", "1-3" or "1,4,7". |
| A clearly two-column paper routes as single-column | Open an issue with the probe detail string from the tool result — it carries the per-page "multi-column share / column count" evidence. A synthetic sample that reproduces the problem is ideal, and tests/fixtures/generate.py is a good starting point. |
Development
No build step and no install step. A bare clone runs the tests immediately:
git clone https://github.com/Dantwz7/dsh-doc-router.git
cd dsh-doc-router
npm test # unit tests run anywhere; integration tests skip without PyMuPDF
npm run test:unit # no external dependencies at all
npm run fixtures # regenerate tests/fixtures/ (needs Python + PyMuPDF)
npm run verify:package # assert what `npm publish` would upload
npm run verify:docs # links, anchors, structure and the README verdict table
lib/index.js the plugin: tool definitions, interpreter discovery, spawning
lib/docprobe.py stdlib + PyMuPDF: format sniffing, column geometry, extraction
cordis.patch.yml bundle layer: inserts the plugin into the DSH composition
tests/ node --test suite + synthetic fixtures
scripts/ fixture generation, package-content and README-fact verification
.github/workflows/ CI (unit and integration both run on Linux and Windows)
One further check needs a real DSH installation:
tests/manual/real-api.mjs drives the plugin with the
real defineTool and schemastery out of the DSH app.asar — the load-time
surface that a cold start fails on.
Routing rules are meant to stay deterministic. If you change a threshold in
lib/docprobe.py, update the rationale in its docstring and the table above, and
add a fixture that exercises the new boundary. See
CONTRIBUTING.md for the non-negotiables and the release process.
Roadmap
- Publish to npm so installation no longer needs a
file:path. — done:dsh-doc-routeris live on the public registry. - Localize the probe's advisory notes. — done:
noteLanguageselects English or Chinese notes, andenbecame the default in0.4.0so the notes match the README npm renders. Both languages are asserted, including that a non-default value actually reaches the probe. - Make the fixtures prove reading order. — done:
two-column-reading-order.pdfcarries per-column markers (L01…L15|R01…R15) and the suite asserts the exact sequence, so an interleaved, swapped, reordered or dropped column now fails the build.three-column.pdf(L01…L15|M01…M15|R01…R15) extends this to a middle column, which a "find the gutter down the middle" implementation misses — and which real input does contain: one paper in the trial corpus is a genuine three-column Science article. - Stop the 25-character block filter from hiding a whole column. — done:
the filter is column-agnostic, so when one column of a two-column page
consisted only of short blocks — a figure, a table, a list of short entries —
it contributed no evidence, the other column looked like a confident
single-column page, and the document was routed to
markitdown: the unsafe direction._column_profilenow reportsunknown(→pdf_markdown) when the discarded blocks lie outside the surviving text span. Safe-side only: it never asserts multi-column from discarded evidence. Re-run over the whole 71-paper trial corpus, 0 verdicts changed, so nothing the thresholds were tuned for regressed;short-column.pdfpins the behaviour. - Count the block filter's threshold in content, not codepoints. — done:
the threshold is a content threshold ("about half a line of body text") but
was counted in raw codepoints, which is not the same amount of text in every
script. A Chinese word is 1–2 ideographs where an English word is ~5 letters,
so the filter was several times more aggressive on Chinese PDFs and discarded
real body text — including, potentially, a whole column of it. It now counts
latin-equivalent codepoints (
CJK_CHAR_WEIGHT = 2), which leaves Latin text byte-identical (0 of 71 corpus verdicts changed) while keeping Chinese body text.dense-script-column.pdfpins both halves: the Chinese column is recovered, and its Latin control still discards. - Recognise UTF-16/UTF-32 plain text. — done: the sniffer only ever
tried UTF-8, so a UTF-16 file fell into
unknown, whose advice ("try it, then use human judgement") is actively misleading for a file that is simply text. It now checks BOMs and, with none, triesutf-8 → utf-16-le → utf-16-be, with a control-character guard so a binary that decodes "successfully" is still rejected.nul-bytes.binis the direct test — it is valid UTF-8 and valid UTF-16LE. - Fail loudly on a broken probe/plugin contract. — done:
doc_routeused to fall back to['markitdown']when the probe returned norecommendat all, dressing a broken contract up as a normal result — in the one direction the asymmetric bet forbids. It now throws with the probe script's path, and a test asserts nothing is returned. - Localize everything the model reads. — done:
noteLanguageused to reach only the probe's advisory notes.pdf_markdown's refusal message and the wholedoc-routingskill body (description,whenToUseand content) now follow it too, so an English session no longer gets a Chinese skill. - Cover the untested format branches and machine-check the README.
— done: 19 new fixtures (tsv / xml / ipynb / md / no-extension /
UTF-16 / UTF-32 / empty / two binaries / jpeg / gif / webp / odt / epub /
mp3) close six format branches, two of which (
odt,epub) had never executed in CI.scripts/check-readme-facts.mjsnow asserts the README's verdict table againstgenerate.py, and the unit suite asserts the skill's tool list matches the probe'srecommendtokens. - Pluggable PDF backend, so a permissively licensed engine can replace PyMuPDF where AGPL is not an option.
- TypeScript declarations for the exported surface.
- Optional OCR route for scanned pages, which currently stop at "render to PNG and look at it".
Contributing
Issues and pull requests are welcome. Bug reports are far more actionable with the
full doc_route result for the file — the
issue template
asks for exactly that. Please read
CONTRIBUTING.md and the
Code of Conduct first.
License
MIT — see LICENSE.
⚠️ PyMuPDF is AGPL-3.0 or commercial
pdf_markdown and PDF layout detection go through PyMuPDF, which Artifex licenses
under the GNU AGPL-3.0 or a commercial license. The distinction matters:
- This plugin's own code is MIT. It does not bundle or link PyMuPDF; it starts an interpreter as a subprocess and exchanges JSON over stdio.
- If you distribute PyMuPDF alongside this plugin (a container image, an installer, a packaged desktop app), your distribution now contains an AGPL-3.0 program, and you must satisfy AGPL-3.0 for it or buy a commercial license from Artifex.
- If you run this as part of a network-accessible service, AGPL §13 may require you to offer the corresponding source to users of that service.
Peer dependencies (@deepseek-ai/dsh-tools BSD-3-Clause; @deepseek-ai/cordis and
@deepseek-ai/schemastery MIT) are supplied by DSH and never bundled. Full details
are in NOTICE. That file describes upstream licensing; it is not legal
advice.
Acknowledgements
- PyMuPDF /
pymupdf4llm— layout-aware PDF extraction, AGPL-3.0 (see above). - MarkItDown — the default converter for everything that is not a multi-column PDF.
- DeepSeek Harness — the host this plugin is written against.
Trademarks belong to their respective owners; this project is not affiliated with or endorsed by any of them.
Related plugins
archify (deepseek-harness)
tt-a1i/archify
WeKnora (dsh-weknora)
tencent/weknora
weknora
tencent/weknora
BrowserSkill (dsh-plugin-browserskill)
tencent/browserskill