Data Platform Engineer with PostgreSQL + Supabase + Ingestion Experience
About the Role
We are looking for an experienced Data Platform Engineer to help professionalize and scale an existing clinical-trial data aggregation platform.
The product is already live as an MVP and uses Supabase/PostgreSQL as its core backend. It collects data from external clinical-trial registries and processes it through several stages, including ingestion, translation, enrichment, combination, and delivery through customer-facing interfaces.
The initial platform was built quickly with the help of AI development tools. The next stage is to bring it to production-grade engineering standards.
You'll begin by developing a deep understanding of the existing database and data pipelines, identifying correctness and reliability issues, fixing known defects, and establishing stronger engineering controls. Longer term, you'll help expand the platform to additional international trial registries.
This is a hands-on role for an experienced engineer who is comfortable inheriting an existing system, investigating unfamiliar data and code, and independently taking problems from diagnosis through production-ready resolution.
What You'll Work On
Audit & Stabilize the Existing Data Platform
-
Independently audit the live Supabase/PostgreSQL database, including its schemas, functions, ingestion processes, and data integrity.
-
Compare findings against known defects and identify additional sources of incorrect, incomplete, stale, or silently missing data.
-
Investigate issues across the complete data lifecycle:
Source Data → Ingestion → Translation → Enrichment → Combination → Database → Dashboard/API
-
Repair broken enrichment fields, missing writes, incorrect filters, and mismatched record counts.
-
Perform safe historical corrections and backfills where required.
-
Identify root causes and implement safeguards that prevent the same classes of failures from recurring.
Build Reliable Data Pipelines
- Migrate existing locally running scrapers and ingestion jobs to hosted infrastructure capable of running unattended on a daily schedule.
- Build ingestion workflows that are idempotent, observable, and safe to retry.
- Implement reconciliation and validation processes to verify that external source records are represented accurately downstream.
- Improve monitoring and failure reporting so missing or incorrect data is detected before it reaches customers.
- Build safe processes for incremental updates and historical backfills.
- Over time, create ingestion pipelines for additional international clinical-trial registries, including sources in Europe, the UK, Japan, and other markets.
Professionalize the Data Infrastructure
- Bring database schemas and functions under version control.
- Establish declarative database migrations and reliable code review/deployment workflows.
- Create clear separation between staging and production environments.
- Establish and test a reliable point-in-time recovery process.
- Improve PostgreSQL performance through indexing, query-plan analysis, and database optimization.
- Introduce engineering controls that make production database changes reproducible, reviewable, and safer.
Data Delivery
- Build and maintain daily full-corpus exports to cloud object storage.
- Develop and maintain a customer-facing REST API.
- Ensure exported and API-delivered datasets remain consistent with the underlying database and original source records.
Required Experience
PostgreSQL & Database Engineering
Strong production experience with PostgreSQL, including:
- Complex schema design and maintenance
- Advanced SQL and PL/pgSQL
- Query-plan analysis and performance optimization
- Indexing strategies
- Transaction boundaries and data consistency
- Schema migrations
- Production database troubleshooting
- Safe data corrections and backfills
We're looking for someone with database engineering depth beyond standard application-level SQL or ORM usage.
Supabase
Hands-on production experience with Supabase is strongly preferred.
Relevant experience includes:
- Supabase CLI
- Declarative migrations
- Database branching or staging workflows
- PostgreSQL functions
- Row Level Security (RLS)
pg_cron
- Deno/TypeScript Edge Functions
You do not necessarily need deep experience with every Supabase feature above, but you should be comfortable working independently within a production Supabase environment.
Data Pipeline Engineering
Experience building and operating production data pipelines, including:
- Incremental ingestion
- Idempotent processing
- Retries and failure recovery
- Data reconciliation and validation
- Safe backfills
- Deduplication and consistency controls
- Monitoring and observability
- Scheduled and unattended ingestion jobs
Experience with Change Data Capture (CDC) or comparable incremental synchronization patterns is valuable.
Data Quality & Reliability
You should have a strong instinct for identifying data-quality problems such as:
- Silent pipeline failures
- Missing or incomplete writes
- Duplicate records
- Incorrect transformations
- Inconsistent counts between systems
- Stale data
- Incorrect aggregations or denominators
- Unverified source references
You should be comfortable tracing an issue from its original external source through ingestion, transformation, storage, and customer-facing delivery.
Programming
Strong proficiency in at least one of:
You should be comfortable building production ingestion workers, integrations, or web scrapers rather than working exclusively with analytics or transformation tools.
Ownership & Working Style
This role requires significant autonomy.
You should be comfortable:
- Inheriting a production system you did not build
- Auditing an AI-assisted MVP and identifying areas that require stronger engineering practices
- Independently investigating unfamiliar data and code
- Identifying root causes rather than only fixing symptoms
- Proposing pragmatic solutions without unnecessarily rebuilding the platform
- Implementing and validating production changes safely
- Clearly documenting findings, risks, and recommendations
Strong written English communication is required for asynchronous remote collaboration, along with a few hours of daily timezone overlap for real-time collaboration with the team.
Nice to Have
- Google Cloud Storage or comparable object-storage experience
- Cloud IAM and service account management
- Web scraping and external-data ingestion
- Experience handling anti-bot protections, proxies, or residential IP infrastructure
- Translation or multilingual data pipelines
- Chinese or Japanese-language data processing
- Pharmaceutical, clinical-trial, healthcare, or life-sciences data experience
Domain experience is helpful but not required. Strong database and data-engineering fundamentals are more important.
What Success Looks Like
The goal is not simply to fix today's bugs. We want to establish a data platform that can be trusted and expanded confidently.
Success means:
- Data pipelines run reliably without manual intervention
- Failures are visible and actionable rather than silent
- Source and destination records can be systematically reconciled
- Production database changes are version-controlled and reproducible
- Historical corrections and backfills can be performed safely
- Database recovery procedures are tested
- Customer-facing data can be trusted
- New international registries can be added through repeatable ingestion patterns
- The platform no longer depends on fragile MVP-era processes to operate reliably