PDF accessibility remediation — Sample: a county document library (140 documents, 1,803 pages)

Machine-checkable PDF/UA-1 defects before and after automated remediation. Every figure on this page comes from veraPDF, the reference open-source validator — not from our own checker.

Documents
140
all processed
Pages
1,803
5611.1s total
Failures before
732,765
veraPDF PDF/UA-1
Failures after
25,526
96.5% resolved; 96.6% excluding OCR documents
Fully conformant
40/140
zero remaining failures
Text recovered
630,797
characters via OCR on 33 doc(s)
Left untouched
0
none made worse
Alt-text cost
$2.72
$0.0015/page

Per document

Before After Bars share one scale; counts are labelled.
DocumentFailed checksResolved / recovered ElementsPer pageConformant
doc92.pdf1 pages · full mode
200,994
2
100.0%8,569149.440sno
doc91.pdf1 pages · full mode
192,609
2
100.0%6,681148.300sno
doc191.pdf385 pages · rebuilt mode
117,795
20,833
82.3%4603.209sno
doc135.pdf1 pages · full mode
27,639
2
100.0%2,45224.800sno
doc139.pdf1 pages · full mode
27,639
2
100.0%2,45224.170sno
doc159.pdf1 pages · full mode
27,639
2
100.0%2,45216.730sno
doc289.pdf206 pages · rebuilt mode
27,310
1,209
95.6%2,1020.477sno
doc290.pdf120 pages · rebuilt mode
19,604
269
98.6%3,4170.987sno
doc154.pdf85 pages · full mode
17,954
2
100.0%5,7600.070sno
doc93.pdf1 pages · full mode
16,736
2
100.0%5289.280sno
doc155.pdf85 pages · full mode
12,807
2
100.0%4,6470.049sno
doc157.pdf64 pages · full mode
9,707
2
100.0%3,5580.033sno
doc153.pdf17 pages · full mode
3,514
2
99.9%1,0910.061sno
doc100.pdf4 pages · rebuilt mode
2,622
115
95.6%3810.207sno
doc168.pdf3 pages · rebuilt mode
2,560
28
98.9%3391.530sno
doc98.pdf4 pages · rebuilt mode
2,319
13
99.4%3161.435sno
doc162.pdf17 pages · metadata mode
2,236
26
98.8%00.115sno
doc156.pdf12 pages · full mode
2,153
24
98.9%6920.168sno
doc167.pdf2 pages · rebuilt mode
1,641
17
99.0%1973.210sno
doc158.pdf6 pages · full mode
952
2
99.8%3460.038sno
doc101.pdf2 pages · rebuilt mode
913
8
99.1%880.290sno
doc190.pdf25 pages · metadata mode
808
211
73.9%00.083sno
doc303.pdf5 pages · full mode
779
4
99.5%1740.060sno
doc102.pdf4 pages · metadata mode
754
27
96.4%00.352sno
doc108.pdf2 pages · full mode
754
4
99.5%2530.545sno
doc128.pdf2 pages · full mode
754
4
99.5%2530.335sno
doc304.pdf5 pages · full mode
745
4
99.5%1760.036sno
doc99.pdf4 pages · metadata mode
621
26
95.8%01.998sno
doc149.pdf8 pages · metadata mode
597
592
0.8%00.060sno
doc117.pdf2 pages · full mode
553
6
98.9%780.200sno
doc130.pdf2 pages · rebuilt mode
486
15
96.9%1112.705sno
doc181.pdf2 pages · rebuilt mode
486
15
96.9%1112.305sno
doc186.pdf1 pages · rebuilt mode
425
2
99.5%710.070sno
doc184.pdf3 pages · full mode · OCR
399
3
+725 chars1705.840sno
doc192.pdf4 pages · full mode
384
3
99.2%520.058sno
doc121.pdf2 pages · rebuilt mode
359
4
98.9%1043.290sno
doc179.pdf2 pages · rebuilt mode
359
4
98.9%1042.510sno
doc123.pdf1 pages · full mode
321
0
100.0%278.110syes
doc182.pdf1 pages · full mode
321
0
100.0%275.420syes
doc137.pdf1 pages · full mode
290
4
98.6%230.100sno

Showing the 40 documents with the most failures. The other 100 went from 5,227 to 2,034 failed checks (0 of them started clean). Every total on this page covers all 140 documents.

What is still failing

RuleDescriptionChecks
7.21.8-1A PDF/UA compliant document shall not contain a reference to the .notdef glyph from any of the text showing operators, regardless of text rendering mode, in any content stream20,683
7.1-3Content shall be marked as Artifact or tagged as real content1,200
7.21.7-1The Font dictionary of all fonts shall define the map of all used character codes to Unicode values, either via a ToUnicode entry, or other mechanisms as defined in ISO 14289-1, 7.21.7857
7.3-1Figure tags shall include an alternative representation or replacement text that represents the contents marked with the Figure tag as noted in ISO 32000-1:2008, 14.7.2, Table 323582
7.21.4.1-1The font programs for all fonts used for rendering within a conforming file shall be embedded within that file, as defined in ISO 32000-1:2008, 9.9479
7.2-10TR element may contain only TH and TD elements369
7.1-1Content marked as Artifact should not be present inside tagged content361
7.1-2Tagged content should not be present inside content marked as Artifact361
7.21.4.2-2If the FontDescriptor dictionary of an embedded CID font contains a CIDSet stream, then it shall identify all CIDs which are present in the font program, regardless of whether a CID in the font is referenced or used by the PDF or not144
5-1The PDF/UA version and conformance level of a file shall be specified using the PDF/UA Identification extension schema100
7.1-5All non-standard structure types shall be mapped to the nearest functionally equivalent standard type, as defined in ISO 32000-1:2008, 14.8.4, in the role map dictionary of the structure tree root. This mapping may be indirect; within the role map a non-standard type can map directly to another non-standard type, but eventually the mapping shall terminate at a standard type82
7.18.5-1Links shall be tagged according to ISO 32000-1:2008, 14.8.4.4.2, Link Element54

Text recovered by OCR

These documents were images with no text layer — unreadable to a screen reader, unsearchable, impossible to copy from. OCR added 630,797 characters of real text across 33 document(s). Expect the failure count to rise on some of them: a blank scan has almost nothing to fail, and readable text has much more. That is an improvement, not a regression. OCR is a transcription and makes mistakes — spot-check anything legally operative.

DocumentCharactersOCR timeFailed checks
doc185.pdf129,04939.9s84 → 81
doc193.pdf129,04970.3s84 → 81
doc287.pdf91,29571.1s53 → 0
doc305.pdf91,29533.6s53 → 0
doc360.pdf54,13525.6s117 → 30
doc366.pdf16,44610.5s14 → 0
doc358.pdf15,15411.1s12 → 0
doc363.pdf13,51814.7s12 → 0
doc351.pdf11,4989.3s12 → 0
doc132.pdf9,6419.3s7 → 0
doc115.pdf7,18411.9s7 → 0
doc364.pdf5,77911.8s7 → 0
doc183.pdf5,6008.7s8 → 0
doc114.pdf4,97713.1s7 → 0
doc120.pdf4,8208.5s7 → 0
doc365.pdf4,3726.8s7 → 0
doc350.pdf4,2609.0s10 → 19
doc343.pdf4,12312.0s11 → 19
doc200.pdf3,3698.6s6 → 0
doc201.pdf3,2248.9s6 → 0

Pages with no text layer

These pages are images of documents. Tagging cannot make them readable, and a validator will not object — a page marked up correctly as a figure passes the structural rules while a screen reader still gets nothing. Any document below needs OCR first; until then, treat a low failure count on it as meaningless.

DocumentScanned pagesShareAction
doc140.pdf1 / 1100%needs OCR before tagging is meaningful
doc160.pdf1 / 1100%needs OCR before tagging is meaningful
doc95.pdf1 / 250%needs OCR before tagging is meaningful
doc234.pdf26 / 7634%needs OCR before tagging is meaningful
doc366.pdf1 / 1010%scanned inserts; OCR those pages
doc190.pdf1 / 254%scanned inserts; OCR those pages
doc141.pdf1 / 303%scanned inserts; OCR those pages
doc289.pdf5 / 2062%scanned inserts; OCR those pages
doc191.pdf7 / 3852%scanned inserts; OCR those pages
doc290.pdf2 / 1202%scanned inserts; OCR those pages

Needs human review

Automated tagging is a first pass. These documents are not ready to publish until a person has confirmed reading order, heading levels, table headers, and every generated alt-text string. Nothing here should be described to an auditor as certified compliant.

…and 2655 more.

Known limitations