arXiv Submission Preparation & Verification

latex_it provides automated packaging and verification for uploading papers to arXiv.org. A single command produces a clean, self-contained submission archive while stripping private comments and verifying that the output compiles identically.


1. Quick Start

To generate an arXiv submission package:

# Package the detected main document
l --arxiv

# Package a specific document
l --arxiv paper.tex

# Package and extract metadata (title, authors, abstract)
l --arxiv --meta paper.tex

This creates:

[!NOTE] For preparing publisher/journal archives (IEEE, Springer, Elsevier) that require flat .tex files without stripping comments or applying arXiv-specific constraints, use -Z / --zip-flat. See Packaging Modes Comparison.


2. Packaging Pipeline (--arxiv)

When --arxiv is invoked, latex_it performs the following steps:

Monolithic TeX Flattening

arXiv prefers a single .tex file or a shallow hierarchy. The flattener (LaTeXFlattener):

Comment & Macro Sanitization

Active Figure Discovery

BibLaTeX Version Shielding

arXiv’s TeX Live environment may run a different version of biblatex than your local machine, leading to wrong format version errors. When biblatex is detected, latex_it:


3. Automated Verification

Before finalizing the package, latex_it tests the generated zip inside an isolated /tmp sandbox:

  1. Sandbox Compilation: Unpacks the archive and compiles it using latex_it --no-env to ensure no ambient environment variables (TEXINPUTS, BIBINPUTS) are required.
  2. Text Layout Verification: Uses pdftotext -layout to compare the rebuilt PDF against the original local PDF. Any text divergence triggers a verification failure.
  3. Visual Page Rendering (pdftoppm): Renders every page at 150 DPI and compares pixel output between builds.
    • Requires pdftoppm (from poppler-utils).
    • Use --no-arxiv-visual-verify to fall back to text-only verification if pdftoppm is not installed.
    • Use --no-arxiv-verify to skip sandbox verification entirely.
  4. Author Verification: Parses expected author names from the TeX source and verifies that each author appears on the first page of the generated PDF. Normalizes accents, umlauts, whitespace, and punctuation. Identifies near-match misspellings via edit distance.

4. Metadata Extraction (--meta)

The --meta option extracts paper metadata and formats it for arXiv’s web submission form:

l --meta paper.tex

Output is saved to arxiv_<document>_meta.txt and contains:


5. Testing & Sampling Tooling

The repository includes tools for sampling and testing real-world arXiv papers:

Sampling Papers (tools/sample_arxiv)

Downloads random source packages from arXiv for testing:

tools/sample_arxiv --output examples/arxiv --attempts 10

Inspecting Metadata (tools/check_arxiv_metadata)

Verifies that the metadata extractor accurately matches arXiv’s API records:

tools/check_arxiv_metadata examples/arxiv/1510.00949v1
tools/check_arxiv_metadata examples/arxiv/1510.00949v1 --main paper.tex --pdf paper.pdf

Full Sandbox Testing (tools/test_arxiv)

Runs an end-to-end audit of a downloaded paper inside a disposable Bubblewrap sandbox:

tools/test_arxiv examples/arxiv/1510.00949v1

Tests include clean builds, incremental caching (--fast), single-pass builds, PDF diffing, portable zip packaging, and arXiv visual verification.

Automated Test Cycle (tools/arxiv_test_cycle)

Combines sampling and testing into a single command:

tools/arxiv_test_cycle --output examples/arxiv