Skip to content

Published results

Explore scientific PDF-to-JSON extraction across four public arXiv papers, with standalone facts, equations, visual findings, and named entities. The corpus spans model architecture, mathematical finance, astrophysics, and language-model experiments: 74 pages in total. Both archived runs used DeepSeek V4.1 Flash through OrcaRouter and completed every phase.

These outputs let you inspect the scientific records available before building a retrieval index or knowledge graph. The use cases explain how to use them in RAG, GraphRAG, and semantic search applications.

Several small and mid-sized open-weight language models were evaluated on this workflow. The aim was to process papers at scale while keeping cost per page low; larger models may improve extraction quality. Only the strongest model on this corpus, DeepSeek V4.1 Flash, is retained in the public showcase.

Results by paper

Model cells contain facts / equations / visuals and link to the full document JSON.

PaperPagesDeepSeek T=0DeepSeek T=1
GQA772 / 4 / 775 / 1 / 7
Dynamic AMM fees23177 / 47 / 23180 / 35 / 24
Stellar XUV flux23326 / 24 / 18240 / 23 / 22
Algospeak21168 / 13 / 21169 / 14 / 19
RunFactsEquationsVisualsCandidatesMainSecondaryCompleted documents
DeepSeek T=074388698723113604
DeepSeek T=166473724001831524

Entities are named scientific methods, systems, or concepts. Candidate counts show what the page pass found; main and secondary counts show what document selection retained. Both runs completed all required phases. Generated content is uncorrected; counts measure output volume, not accuracy or recall. Parser warnings remain visible in the raw request artifacts.

DeepSeek V4.1 Flash at temperature 0 preserved more equations and produced fewer document-dependent statements on this corpus. This observation is specific to the measured model and corpus.

The outputs still require entity normalization, cross-page deduplication, and citation-graph construction before large-scale indexing.

Raw outputs

Each run archive contains its documents, prompts, complete provider responses, checkpoints, GROBID TEI, configuration, and request statistics.

See cost and timing, reproduction instructions, and source attribution.

Released under the Apache License 2.0.