Outlink AI is building automated outreach software for link building — the platform finds relevant, high-authority sites, writes personalized pitches, sends them from a customer's own inbox, handles follow-ups and replies, and verifies live placements, all inside guardrails the customer sets. We're looking for a Data Engineer to help build and scale the data infrastructure that powers prospect discovery, qualification, and placement verification across the product.
What you'll do
- Design, build, and maintain data pipelines that support prospect discovery, relevance/authority scoring, and spam/PBN filtering at scale
- Develop and optimize ETL/ELT processes to ingest data from web crawls, third-party APIs, and internal product events
- Build and maintain the data models and warehouse structures that power analytics, reporting, and product features like the weekly digest
- Ensure data quality, consistency, and freshness across pipelines feeding real-time scoring and verification systems
- Collaborate with backend engineers and data scientists to productionize models used for relevance matching and reply classification
- Monitor pipeline performance, troubleshoot data issues, and improve system reliability and scalability as data volume grows
What we're looking for
- 3+ years of experience as a data engineer or in a similar backend/data-focused engineering role
- Strong proficiency in SQL and at least one programming language commonly used for data engineering (e.g., Python, Scala, or Java)
- Experience building and maintaining ETL/ELT pipelines using modern orchestration tools (e.g., Airflow, Dagster, or similar)
- Familiarity with cloud data warehouses (e.g., Snowflake, BigQuery, Redshift) and data pipeline tooling
- Solid understanding of data modeling, schema design, and performance optimization for large datasets
- Comfortable working in a fast-moving environment with evolving priorities and close collaboration with product and engineering teams
Nice to have
- Experience working with web crawling, scraping, or large-scale text data
- Familiarity with email deliverability, domain reputation, or similar adjacent domains
- Experience supporting machine learning or NLP-based classification systems in production
