Synthetic Onboarding — Equal Employment Opportunity Data Data

Synthetic training data — no real PII, fully coherent identities

HR

Generate synthetic EEO self-identification forms — voluntary gender and race or ethnicity checkboxes feeding EEO-1 reporting, with the severe class imbalance real demographic data actually has.

14

Fields per document

1

Page

HR

Category

What this document is

The equal employment opportunity self-identification form asks a new hire to voluntarily report gender and race or ethnicity, alongside their job title and work location, and to sign. Employers aggregate the responses into EEO-1 reporting to the federal government. It is voluntary, it is separated from the hiring file, and it is the source of every workforce demographic number an employer publishes.

Why generate synthetically

Demographic self-identification data is categorically unshareable — it is the most protected category of employment record there is, kept apart from the personnel file by design. Yet the forms have to be read, aggregated and reported accurately, and a misread category corrupts a federal filing. The other problem is statistical: real demographic distributions are severely imbalanced, and a corpus that balances them to make training easier produces a model that has never seen how rare the rare categories actually are.

What makes synthetic data useful

The distributions here are deliberately not balanced. Gender splits 51.5% female to 48.5% male; race or ethnicity runs 74.7% white, 13.3% Black, 5.9% Asian, 5.6% two or more races and 0.3% Native American, with the Hispanic or Latino ethnicity question marked on 22.5% of forms as the separate two-part question the EEO-1 instrument actually uses. Every form carries a job title, a work location, a signature, a printed name and a date, so the demographic marks arrive attached to a realistic employment record.

Training challenges

The rarest category appears on two forms in a thousand-identity corpus. That is the honest prevalence, and it means a detector trained here will effectively never have seen that box marked — which is the correct lesson about the problem, not a defect in the data. It also means accuracy is a useless metric on this form: a model that never marks anything but the modal category scores above 74%. The two-part ethnicity-then-race structure is a second trap, because the Hispanic or Latino question is a separate mark rather than one of the race options, and models routinely collapse them into a single group.

Generate synthetic Onboarding — Equal Employment Opportunity Data data

Start with 500 free credits. No credit card required.

Generate Now

Who uses this data

HR-tech onboarding suites capturing self-identification, HRIS and compliance vendors aggregating EEO-1 reporting from scanned forms, workforce-analytics platforms, and applicant-tracking providers that must keep demographic responses segregated from the hiring record while still processing them accurately.

Document complexity profile

14 fields on a single page: 4 text, 8 checkbox targets across a gender group, a separate ethnicity mark and a five-way race group, plus a signature and a date, with one label-to-value relation per field. 1 arithmetic binding and 2 function calls, at maximum expression depth 2. The checkbox distribution is severely imbalanced by design.

Key stats from our synthetic corpus

Quantitative characteristics of the Onboarding — Equal Employment Opportunity Data documents our generator produces.

MetricValueDetail
Rarest category prevalence0.3%Native American is marked on 0.3% of forms — two documents in a thousand-identity corpus. That is the honest prevalence, and it is the reason accuracy is the wrong metric for this document.
Modal category share74.7%74.7% of forms mark white, 13.3% Black, 5.9% Asian and 5.6% two or more races. A model that always predicts the mode scores above 74% and reports nothing useful, so score per-category recall instead.
Ethnicity question marked22.5%The Hispanic or Latino question is a separate mark preceding the race question, present on 22.5% of forms. Collapsing it into the race group is the most common structural error on this document.
Gender split51.5% / 48.5%Female 51.5%, male 48.5%, as a mutually exclusive pair. The one near-balanced group on the page, and a useful control against the severely imbalanced ones beside it.
Signature present100%Every form carries a signature, a printed name and a date alongside job title and work location, so the demographic marks always arrive attached to an employment record rather than floating alone.

How this document co-occurs with others

Rates at which identities in our corpus that produce a Onboarding — Equal Employment Opportunity Data also produce other documents.

CorrelationRateDetail
Employee record in the same packet100%The personal information sheet carries the same employee's name and job title. In a real HR file these two are deliberately stored apart, which makes packet-level routing a compliance requirement and not just a convenience.
Employment eligibility verification100%Completed in the same packet. The I-9 records citizenship status, which is a different legal category from the ethnicity data here and must not be conflated with it by a downstream system.
Other short signed page100%Both are sparse signed attestations arriving in the same upload — the pair that classifiers keying on marks rather than layout most often swap.
Demographic capture in a different system100%The healthcare intake form asks the same person for ethnicity in a different structure and vocabulary. Comparing the two exposes whether a model has learned an instrument or a concept.

Demographic distributions above are corpus-derived: they were computed over the 641 employment-eligible identities inside a local synthetic corpus of 1,000 identities generated by SymageDocs' World Simulation Engine at seed 20260421 — the shipped definition gates onboarding forms on active employment. Field, type and relation counts come from the shipped equal employment opportunity definition in the SymageDocs form library. No real employee or demographic data was used at any stage.

Frequently asked questions

What data format do synthetic EEO self-identification forms include?
Each generated identity produces a rendered PDF plus a structured JSON annotation file with bounding boxes, field types, and ground-truth values for all 14 fields — 4 text, 8 checkbox targets, a signature and a date — with one label-to-value relation per field. COCO, YOLO, FUNSD, and BIO/NER exports come from the same job.
Why is the class distribution so imbalanced?
Because real workforce demographics are. Rebalancing the categories would make training easier and evaluation meaningless: the point of this form is that a production model has to handle a category that appears on a fraction of a percent of documents without collapsing to the majority class.
How is the Hispanic or Latino question handled?
As a separate mark, the way the EEO-1 instrument structures it — an ethnicity question asked before and independently of the race question. It is marked on 22.5% of forms. Treating it as a seventh race option is the most common modelling error on this document.
What metric should I use on this form?
Not accuracy. A model that always predicts the modal category scores above 74% while being useless for the reporting the form exists to support. Per-category recall, or macro-averaged F1, is the metric that reflects whether the form is being read.
Can I use this data commercially?
Yes. Every employee, job title, location and demographic response is synthetic, contains no real personal or employment data, and is licensed for commercial use including model training, benchmarking, and redistribution inside your own products.

Related HR Forms