Casework produces conformity files, Fundamental Rights Impact Assessments, and Article 9 risk documentation for AI hiring tools, aligned to the EU AI Act, Colorado AI Act, Illinois HB 3773, and NYC LL 144. As the company scales its productized documentation engagements, we're looking for a Data Engineer to help build the data infrastructure and pipelines that power discovery, drafting, and citation-verification work across engagements.
What you'll do
- Design, build, and maintain data pipelines that ingest, transform, and structure information used in regulatory documentation production
- Develop and maintain data models supporting discovery synthesis, regulation mapping, and draft assembly workflows
- Build tooling to verify data quality, lineage, and citation accuracy across generated artifacts
- Partner with engineering and subject-matter staff to expose clean, reliable data to internal production tools
- Monitor pipeline performance and reliability, troubleshooting issues as they arise
- Evaluate and integrate new data sources as the methodology and vendor landscape evolve
What we're looking for
- Experience building and maintaining production data pipelines (batch and/or streaming)
- Strong SQL skills and experience with at least one modern data warehouse (e.g., Snowflake, BigQuery, Redshift)
- Proficiency in Python or a similar language for data engineering work
- Experience with workflow orchestration tools (e.g., Airflow, Dagster, Prefect)
- Solid understanding of data modeling, schema design, and data quality practices
- Comfort working in a small, fast-moving team where scope and priorities shift as the company grows
Nice to have
- Experience working with document-heavy or unstructured text data
- Familiarity with retrieval or search infrastructure supporting AI-assisted content generation
- Prior experience at an early-stage or productized-services company