Free Synthetic Document Dataset
Download 10 coherent synthetic identities filling 3 common tax and healthcare form layouts — 30 filled documents with FUNSD, BIO, YOLO and COCO annotations, ground-truth JSON, and train/val/test splits, in one ZIP. Fully synthetic: generated algorithmically, not derived from any real individual's records.
No email, signup, or credit card required.
10
Coherent identities
30
Filled documents
140
Page images at 300 DPI
4
Annotation formats
What's in the pack
Each of the 10 synthetic identities fills every form in the pack, so names, employers, wages, and diagnoses stay consistent across a person's documents — the same cross-form coherence the full product generates at scale. The forms are synthetic recreations of common US form layouts, filled with generated data; they are not official documents and are not affiliated with or endorsed by any government agency.
- W-2 Wage and Tax Statement (11 pages)
- Form 1040 - U.S. Individual Income Tax Return (2 pages)
- CMS-1500 Health Insurance Claim Form (1 page)
The ZIP contains: Typed PDFs, Page-image PNGs (typed), BIO labels, YOLO labels, COCO annotations, FUNSD JSON — plus per-document ground-truth JSON, a tabular ground-truth CSV, deterministic train/val/test splits, a machine-readable manifest, and a dataset-card README.
Annotation formats
- FUNSD
- Per-page JSON in the canonical FUNSD shape (form / words / linking) for form-understanding models such as LayoutLMv3 and LiLT.
- BIO
- Token-level BIO tag sequences for NER and token-classification training.
- YOLO
- Normalized bounding-box label files for object-detection training on page images.
- COCO
- Dataset-level COCO annotation JSON with per-page images and box annotations.
Fully synthetic data
Every value in the pack is generated algorithmically. The dataset is intended to be fully synthetic and was not derived from any real individual's records, so there is no real PII to leak into your training pipeline. Because values are generated programmatically, records may coincidentally resemble real people or entities; validate the dataset's suitability for your use.
License
The pack ships under the SymageDocs Sample Data License (symagedocs-sample-pack-1.0, the LICENSE file in the ZIP). In short: free to use for evaluating SymageDocs and for training, fine-tuning, and benchmarking machine-learning models — including commercial models. Redistributing or reselling the pack, and using it to build a competing dataset product, are not permitted. If you publish benchmarks derived from it, credit SymageDocs with a link. The LICENSE file contains the controlling terms.
Get updates when the pack changes
Optional — the download above works without it. Leave an email and we'll let you know when new forms, formats, or fixes land in the Sample Pack.
Need more than 10 identities?
The full product generates datasets at scale across the whole form catalog, adds handwritten surfaces, scan/photo degradation, and Donut metadata output, and starts with free monthly credits.
Create a Free Account