Synthetic W-2 Wage and Tax Statement Data
Synthetic training data — no real PII, fully coherent identities
Generate synthetic 2024 vertical-stack W-2 sheets — four full-width copies stacked down one page at identical horizontal offsets. The layout that strips horizontal position out of the copy-disambiguation problem, with the 2024 Social Security wage base.
144
Fields per document
1
Page
Tax
Category
What this document is
This is the 2024 vertical-stack W-2: one page carrying four full-width Wage and Tax Statements for the same employee, printed one above the other. It is the standard output of laser and continuous-feed payroll printing, and the format in which most employees of mid-size employers receive their statement. Structurally it is the same document as the four-up quadrant sheet; spatially it is the opposite arrangement, and that difference is why the two behave so differently under extraction.
Why generate synthetically
If you want to know whether a document model has learned the W-2 or has learned one W-2 picture, the vertical stack is the experiment. It holds the content, the field set, the wage distribution and the annotation schema constant against the quadrant sheet and changes only the spatial arrangement — and it changes it in the one direction that removes the strongest signal most layout models rely on. Constructing that ablation from collected documents would require finding two employers who print the same year's statement in different formats and getting both to release their files. Generating it takes one job.
What makes synthetic data useful
Each sheet replicates one simulated employee's 2024 statement in four full-width bands with identical values. Box 3 Social Security wages cap at the 2024 wage base of $168,600 while Box 5 Medicare wages continue to the top of the income distribution, and Box 17 state tax is zero for residents of states with no income tax rather than randomly small. The Box 12 codes follow the realistic frequency of elective deferrals, health coverage cost and Roth contributions, replicated identically across all four bands so the corpus supervises copy assignment rather than only value extraction.
Training challenges
Four copies at identical horizontal offsets means the horizontal coordinate of a label predicts nothing about which statement it belongs to. Every association decision has to be made vertically, across bands that are a quarter of a page tall and separated by a printed cut line rather than by whitespace. Because the four copies carry identical values, a model that assigns Box 1 from band one to Box 2 from band three produces output that is numerically correct and structurally wrong — an error class that only copy-level annotation can expose. The wide, short box geometry of a full-width band is also a poor match for detectors pretrained on roughly square regions.
Generate synthetic W-2 Wage and Tax Statement data
Start with 500 free credits. No credit card required.
Generate NowWho uses this data
Income-verification and mortgage platforms that receive whatever their borrower's employer happened to print, payroll bureaus validating their own laser print output, and document-AI teams running layout-robustness evaluations rather than single-format accuracy claims. Benefits and public-assistance eligibility systems see this format constantly because it is what mid-size employers issue.
Document complexity profile
144 fields on a single page: 72 currency amounts, 52 text, 12 checkbox targets, and 4 Social Security number and 4 employer identification number typed fields, arranged as one 36-field statement stacked four times. 108 annotation relations, 12 conditional bindings, 16 function calls, maximum expression depth 2. Field geometry is wide and short — the opposite aspect ratio to the quadrant sheet's fields — at identical horizontal offsets across all four bands.
Key stats from our synthetic corpus
Quantitative characteristics of the W-2 Wage and Tax Statement documents our generator produces.
| Metric | Value | Detail |
|---|---|---|
| 2024 Social Security wage base | $168,600 | Box 3 Social Security wages cap at $168,600 while Box 5 Medicare wages continue to $316,900. The middle rung of the three-year stacked ladder, and the value most current production validators were built against. |
| Horizontal offsets shared across copies | 4 of 4 | All four bands place every field at the same horizontal position. Horizontal coordinate therefore carries zero information about copy identity, which is what makes this layout the sharpest available probe of layout dependence in a W-2 model. |
| Box 12 code distribution | D 72% / E 18% / G 10% | Box 12a codes are 71% D, 17% E and 12% G, populated on 57.1% of sheets; Box 12b is 83% DD and 17% C on 49.8%. Each code appears four times per sheet at four vertical offsets, so a code-classification model gets four independent looks at the same label — and four chances to disagree with itself. |
| Blank field rate | 42% | 60 of the annotated fields render blank on a typical sheet, repeated across the four bands. Every one is annotated, so an extractor that hallucinates plausible values into empty boxes can be caught rather than credited. |
| Median Box 1 wages | $45,600 | Across 641 eligible synthetic identities the median Box 1 wage is $45,600, p25–p75 $27,200 to $74,400. Held identical to the 2023 and 2025 stacks so that cross-revision comparisons are controlled. |
How this document co-occurs with others
Rates at which identities in our corpus that produce a W-2 Wage and Tax Statement also produce other documents.
| Correlation | Rate | Detail |
|---|---|---|
| Same year, quadrant arrangement | 100% | The 2024 four-up sheet carries identical content in a two-by-two grid. Same year, same wages, same annotation schema, opposite spatial arrangement — the paired evaluation that measures layout dependence directly. |
| Next revision of the same stack | 100% | The 2025 stack raises the Social Security cap and ships as a two-page template with a field-free second page. Pairing them adds a page-handling variable on top of the year variable. |
| Withholding election on file | 100% | The same employees have W-4s with their employer. Pairing the election with the resulting Box 2 federal withholding is how payroll and onboarding systems verify that an election was actually applied. |
| Employer's annual unemployment return | 100% | The issuing employer also files Form 940. Aggregating employee statements against the employer's FUTA return is the reconciliation payroll auditors run at year end. |
Wage and rate figures above are corpus-derived: they were computed over the 641 employment-eligible identities inside a local synthetic corpus of 1,000 identities generated by SymageDocs' World Simulation Engine at seed 20260421. Field, type and relation counts come from the shipped 2024 vertical-stack W-2 definition in the SymageDocs form library. No real employee, employer, or payroll data was used at any stage.
Frequently asked questions
- What data format do synthetic 2024 vertical-stack documents include?
- Each generated identity produces a rendered PDF plus a structured JSON annotation file with bounding boxes, field types, and ground-truth values for all 144 fields on the sheet — 36 fields per copy across four stacked copies — plus 108 label-to-value relations, each tagged with the band it belongs to. COCO, YOLO, FUNSD, and BIO/NER exports come from the same job.
- Why is the vertical stack harder than the quadrant sheet?
- Because it removes horizontal position as a disambiguating feature. On a two-by-two grid, a Box 1 label on the left half and a Box 1 label on the right half are trivially distinguishable by x-coordinate. On a vertical stack all four Box 1 labels share the same x-coordinate and differ only in y. Models that lean on horizontal layout — which is most of them, because most wide forms reward it — degrade measurably here.
- Can I compare this directly against the quadrant sheet?
- Yes, and that is the intended use. The same simulated employees generate both layouts with the same wage values and the same annotation schema, so a paired evaluation isolates spatial arrangement as the only variable. Any accuracy gap between the two is a measurement of layout dependence, not of sample difficulty.
- What does Box 1 represent on this layout?
- Gross annual wages. On the stacked and four-up layouts Box 1 carries the employee's gross annual wage, so Box 1, Box 3 and Box 5 agree below the Social Security cap. The half-page single-copy W-2 binds Box 1 to federal wages net of elective deferrals instead; use that layout if your evaluation depends on the Box 1 versus Box 12 relationship.
- Can I use this data commercially?
- Yes. Every Social Security number, name, address, employer and wage figure is synthetic, contains no real personal data, and is licensed for commercial use including model training, benchmarking, and redistribution inside your own products.