All Our TeX Source Are
Not Belong to You

PDF Obfuscation against
arXiv's TeX Source Policies

Zhongtang Luo, Jianting Zhang, Zheng Zhong

Purdue University

$\LaTeX$
$\text{Lua}\LaTeX$

How does arXiv (mis)detect (Lua)LaTeX files?


$ mutool show test.pdf 509

509 0 obj
<<
  /Author ()
  /Title ()
  /Subject ()
  /Creator (LaTeX with hyperref)
  /Keywords ()
  /Producer (LuaTeX-1.21.0)
  /CreationDate (D:19791231190000-05'00')
  /ModDate (D:19791231190000-05'00')
  /Trapped /False
  /PTEX.FullBanner (This is LuaHBTeX, Version 1.21.0 \(TeX Live 2025/nixos.org\))
>>
endobj
                

$ mutool show test.pdf 509

509 0 obj
<<
  /Author ()
  /Title ()
  /Subject ()
  /Creator (LaTeX with hyperref)
  /Keywords ()
  /Producer (LuaTeX-1.21.0)
  /CreationDate (D:19791231190000-05'00')
  /ModDate (D:19791231190000-05'00')
  /Trapped /False
  /PTEX.FullBanner (This is LuaHBTeX, Version 1.21.0 \(TeX Live 2025/nixos.org\))
>>
endobj
                

$ mutool show test.pdf 509

509 0 obj
<<
  /Author ()
  /Title ()
  /Subject ()
  /Creator (LaTeX with hyperref)
  /Keywords ()
  /Producer (LuaTeX-1.21.0)
  /CreationDate (D:19791231190000-05'00')
  /ModDate (D:19791231190000-05'00')
  /Trapped /False
  /PTEX.FullBanner (This is LuaHBTeX, Version 1.21.0 \(TeX Live 2025/nixos.org\))
>>
endobj
                

tex?

latex?

text?

texas?

Field Data Undetected?
/Creator TeX
/Creator ConTeXt
/Creator texture
/Creator plaintext
/Creator pdfTeXt
/Creator texttext
/Creator texttex
/Creator tex-text
/Creator tex
/Creator texttex
/Creator ConTeX
/Creator TeX*
/Creator *TeX
/Creator LuaLaTeX
/Creator pdfTeXs
/Creator texas
/Creator TeXas
/Creator vertex
/Creator paleocoretex
Field Data Undetected?
/Producer pdfTeX-1.40.26
/Producer pdfConTeXt-1.40.26
/Producer texture
/Producer plaintext
/Producer pdfTeXt
/Producer texttext
/Producer texttex
/Producer tex-text
/Producer pdftex-1.40.26
/Producer pdfConTeX-1.40.26
/Producer pdfTeX*-1.40.26
/Producer pdf*TeX-1.40.26
/Producer pdfLuaLaTeX-1.40.26
/Producer TeX
/Producer pdfTeXs
/Producer texas
/Producer TeXas
/Producer vertex
/Producer paleocoretex
Field Data Undetected?
/Author tex
/CreationDate tex
/ModDate tex
/PTEX.Fullbanner tex
/Subject tex
/Title tex
/Trapped tex

# --- Strip Document Info ---
print(f"[*] Metadata: {pdf.docinfo}")
print("[*] Stripping metadata...")
for key in list(pdf.docinfo.keys()):
    del pdf.docinfo[key]

# --- Remove XMP Metadata if present ---
if "/Metadata" in pdf.Root:
    del pdf.Root["/Metadata"]
print("[*] Removed XMP metadata stream")

$ pdffonts test.pdf
name                                 type              encoding         emb sub uni object ID
------------------------------------ ----------------- ---------------- --- --- --- ---------
ENBKCE+LMRoman10-Regular             CID Type 0C       Identity-H       yes yes yes      4  0
YCHXXL+CMMI10                        Type 1            Builtin          yes yes no       5  0
CZKGAS+CMR10                         Type 1            Builtin          yes yes no       6  0
YTYGMW+CMEX10                        Type 1            Builtin          yes yes no       7  0
FXXUVH+CMSY10                        Type 1            Builtin          yes yes no       8  0
EJASTL+CMR7                          Type 1            Builtin          yes yes no       9  0
                
Field Data Undetected?
Font's FontName QMHNZG+CMR10
Font's FamilyName Computer Modern
Font's FullName CMR10
/FontName /QMHNZG+CMR10
/BaseFont /QMHNZG+CMR10
/BaseFont /QMHNZG+CMMI10
/BaseFont /QMHNZG+CMBX10

with open(renamed_font_path, "rb") as f:
    new_font_data = f.read()
new_stream = pikepdf.Stream(pdf, new_font_data)

new_fontname = safe_fontname(fontname, replacement)
fontdesc[fontfile_key] = new_stream
fontdesc["/FontName"] = pikepdf.Name(new_fontname)

if "/BaseFont" in obj:
    print(f"[+] Updating BaseFont from {obj['/BaseFont']}")
    base_font = str(obj["/BaseFont"])
    if "CM" in base_font:
        new_base = safe_fontname(base_font, replacement)
        obj["/BaseFont"] = pikepdf.Name(new_base)

print(f"[✓] Replaced: {fontname} → {new_fontname}")
                

This works!

But is TeX source detection an achievable goal?


Catalog (13 0 obj)
└── Pages (5 0 obj)
    └── Page (2 0 obj)
        ├── Contents (3 0 obj)
        │   └── "BT /F31 9.96264 Tf 1 0 0 1 140.944 656.037 Tm [<00350049004A0054>-333<004A0054>-333<0042>-333<0055004600540055>-333<005100420053004200480053004200510049000F>]TJ 1 0 0 1 303.509 96.112 Tm [<0012>]TJ ET"
        └── Resources
            └── Font /F31 (4 0 obj, /CNSUIX+NewCM10-Book)

Trailer
└── Info (14 0 obj)
                        

Catalog (18 0 obj)
└── Pages (9 0 obj)
    └── Page (2 0 obj)
        ├── Contents (10 0 obj)
        │   └── "1 0 0 -1 0 841.8898 cm /d65gray cs 0 scn /F0 10 Tf BT 1 0 0 -1 126 132.83 Tm [(\000\001\000\002\000\003\000\004\000\005\000\003\000\004\000\005\000\006\000\005\000\007\000\b\000\004\000\007\000\005\000\t\000\006\000\n\000\006\000\013\000\n\000\006\000\t\000\002\000\f)] TJ ET"
        └── Resources
            └── Font /F0 (4 0 obj, /CMINVQ+NewCM10-Regular-Identity-H)

Trailer
└── Info (16 0 obj)
                        

We build a PDF obfuscation tool that translates the PDF content stream on the left (TeX-generated) to the one on the right (Typst-generated)

Existing detection methods are weak

No intrinsic differences between TeX-generated and Typst-generated PDFs

Goal of distinguishing TeX-generated PDF files from other typesetting systems is fundamentally unsound!

Responsible Disclosure

Hey do you guys think
this is a security vulnerability?
It's OK. Authors can submit PDFs.
arXiv Scientific Director, Prof. S.S.:
Please submit laTeX (sic) source
I am not using LaTeX.
Why do you think so?
Hello?

Exploiting PDF Obfuscation in LLMs, arXiv, and More