Synthetic WH-380-E Certification of Health Care Provider for Employee's Serious Health Condition Data
Synthetic training data — no real PII, fully coherent identities
Generate synthetic WH-380-E FMLA medical certifications for an employee's own serious health condition, with a coherent employee, employer, treating provider, and clinical narrative. Each certification pairs a simulated worker and job description with a provider specialty and condition profile that actually fit each other.
76
Fields per document
4
Pages
HR
Category
What this document is
Form WH-380-E is the U.S. Department of Labor's model certification of a health care provider for an employee's own serious health condition under the Family and Medical Leave Act. The employer completes the job-description block, the employee identifies themselves, and the treating provider fills out four pages of clinical questions: when the condition began, how long it is expected to last, whether it is chronic or an episode of incapacity plus continuing treatment, which essential job functions the employee cannot perform, and how often and for how long incapacity is expected to recur.
Why generate synthetically
An FMLA certification is a medical record about a named employee sitting in an HR file, which puts it simultaneously under employment-privacy expectations and, in practice, HIPAA-adjacent handling rules. Leave administrators process them at enormous volume and almost none of them can be shared outside the organisation, so the leave-management and third-party-administrator market has been building extraction models on synthetic-by-hand samples for years. Generating them removes the sourcing problem and, more usefully, lets you control the mix of chronic versus episodic conditions rather than inheriting whatever mix one employer happened to produce.
What makes synthetic data useful
Each certification is internally consistent across all four pages. The employee's job title and work schedule come from their simulated employment record, and the essential-functions text describes that job rather than a generic one. The treating provider's specialty is drawn to fit the condition — orthopedics for a post-surgical knee, pulmonology for an inhaler-managed respiratory condition, psychiatry for a condition under active medication management — and the medical-facts narrative, the incapacity frequency, the incapacity duration, and the list of functions the employee cannot perform all describe the same underlying condition. The employer's return-by date is set relative to the certification date the way the regulation requires.
Training challenges
The form is a checkbox-heavy clinical questionnaire in which most checkboxes are unchecked on any given document, and the unchecked ones are not blank space — they are printed empty squares with adjacent question text that models routinely read as answered. Questions 5 through 10 are built as "was / is / will be" triplets whose members sit on consecutive lines with nearly identical labels, so the correct answer depends on which of three near-identical rows carries the mark. The free-text clinical answers are multi-line and wrap unpredictably, and the same question can be answered in one line or five, which moves every field below it. Four pages each repeat the employee's name in a header field, giving a document-linking signal that is easy to exploit and easy to get wrong when pages are scanned out of order.
Generate synthetic WH-380-E Certification of Health Care Provider for Employee's Serious Health Condition data
Start with 500 free credits. No credit card required.
Generate NowWho uses this data
FMLA and absence-management administrators, leave-management platforms, and the third-party administrators who process certifications on behalf of large employers — the buyers who receive these forms by fax and scan at volume and need them classified, extracted, and checked for completeness before a leave decision is made. Also relevant to disability and workers' compensation claim intake, which handles structurally similar provider certifications.
Document complexity profile
76 fields spread 11/30/29/6 across the four pages — 41 text and 35 checkbox targets — with 109 annotation relations tying printed question text to the answer field it governs. Binding logic is light (3 conditional bindings, 9 function calls, maximum expression depth 2) because the clinical coherence is produced by a registered computed-field module before evaluation. The difficulty is that checkboxes outnumber text answers on the certification pages and most of them are correctly empty.
Key stats from our synthetic corpus
Quantitative characteristics of the WH-380-E Certification of Health Care Provider for Employee's Serious Health Condition documents our generator produces.
| Metric | Value | Detail |
|---|---|---|
| Chronic versus episodic condition | 77% chronic | 76.6% of generated certifications mark the condition as chronic and 23.4% as an episode of incapacity plus continuing treatment. These two boxes decide whether the leave is administered as intermittent or continuous, so the split is the single most consequential binary on the form. |
| Checkbox targets that render blank | 27 of 35 | 27 of the 35 checkbox targets are correctly empty on every certification — the inpatient-care block, the pregnancy block, and the Question 9 and Question 10 frequency grids. Empty printed squares next to question text are the most common source of false positives in clinical-form extraction, and this corpus supervises them explicitly. |
| Provider specialties represented | 6 | Certifying providers are drawn from six specialties — orthopedics 24%, internal medicine 22%, rheumatology 14%, neurology 14%, pulmonology 13%, and psychiatry 13% — each paired with the conditions that specialty would actually treat. |
| Distinct clinical narratives | 8 | Eight distinct medical-fact narratives circulate across the corpus, each with its own matching incapacity frequency, incapacity duration, and list of essential functions the employee cannot perform. Free-text clinical prose of varying length is what stresses a layout model's handling of reflowing multi-line answers. |
| Work-schedule variants | 6 | Employee work schedules span six patterns including rotating 12-hour shifts and 37.5-hour weeks, not just Monday-to-Friday nine-to-five. Schedule text feeds the intermittent-leave calculation an administrator performs from this form. |
How this document co-occurs with others
Rates at which identities in our corpus that produce a WH-380-E Certification of Health Care Provider for Employee's Serious Health Condition also produce other documents.
| Correlation | Rate | Detail |
|---|---|---|
| Employer's eligibility notice for the same leave | 100% | Every WH-380-E-eligible identity also generates a WH-381, the notice of eligibility and rights the employer must issue within five business days of a leave request. The certification and the notice are two halves of one leave file and arrive in the same scanned batch. |
| Family-member certification, same employee | 100% | The same employee can also file a WH-380-F for a family member. Generating both produces the E-versus-F contrast pair that FMLA document classifiers most often get wrong, because the two forms share a header, a layout family, and most of their label vocabulary. |
| Same employee's onboarding file | 100% | The employee on the certification has an I-9 on file with the same employer. Linking a mid-employment leave document back to the onboarding record is the identity-resolution task an HR document repository has to solve across form types. |
| Employees with partnership income | 13% | 13.4% of FMLA-eligible identities also qualify for a Form 1065 partnership return — employees with outside business interests. It is a reminder that an HR corpus and a tax corpus overlap on the same people, which matters if you are deduplicating identities across document types. |
Rate figures above are corpus-derived: they were computed over the 641 employment-eligible identities inside a local synthetic corpus of 1,000 identities generated by SymageDocs' World Simulation Engine at seed 20260421. Field, page, type and relation counts come from the shipped WH-380-E definition in the SymageDocs form library. No real employee, patient, or provider data was used at any stage.
Frequently asked questions
- What data format do synthetic WH-380-E documents include?
- Each generated identity produces a rendered PDF plus a structured JSON annotation file with bounding boxes, field types, and ground-truth values for all 76 fields across the four pages, plus 109 label-to-value relations linking the printed question text to the answer it governs. COCO, YOLO, FUNSD, and BIO/NER exports come from the same job.
- Does the clinical content actually hang together?
- Yes. The provider specialty, the medical-facts narrative, the expected duration, the incapacity frequency and duration, and the essential functions the employee cannot perform are drawn together as a single condition profile. A certification about lumbar radiculopathy is signed by a specialist who would treat it, describes physical limitations that follow from it, and reports a recurrence pattern consistent with it — so the document supports condition-classification and consistency-checking tasks, not just box-level extraction.
- How does this differ from the WH-380-F you also publish?
- WH-380-E certifies the employee's own serious health condition; WH-380-F certifies a family member's, and adds the family-member identification and care-description blocks while dropping the employee-specific job-function analysis. The two are visually similar and functionally different, which makes them the best available adversarial pair for FMLA document classifiers — a model that routes an F to an E workflow sends a caregiver's certification down the wrong compliance path.
- Which regions vary and which are fixed?
- The identity, employer, job, provider, condition and narrative regions vary across the corpus. The Question 5 through Question 8 treatment block currently resolves to a fixed set of selections and a fixed treatment description, so that region exercises layout, label geometry and blank-versus-marked discrimination rather than answer-class balance. If your evaluation depends on variation inside that block specifically, score it separately.
- Can I use this data commercially?
- Yes. Every employee, employer, provider, and clinical detail is synthetic, describes no real person or patient, and is licensed for commercial use including model training, benchmarking, and redistribution inside your own products.