Research Engineer, Privacy and Anonymization
clera · San Francisco
Job description
About the role
This role focuses on building privacy and anonymization systems that keep sensitive real-world data safe and useful for AI training. You will own the end-to-end pipeline, bridging applied research and production engineering in a fast-moving AI infrastructure company.
Key responsibilities
- Design and implement detection systems for PII, quasi-identifiers, credentials, and other sensitive content, choosing transformations based on data type and downstream use case.
- Develop and benchmark detection approaches that combine rule-based methods, statistical models, classifiers, and LLM-based techniques.
- Build production pipelines that anonymize raw data before it enters training, evaluation, or synthetic data generation workflows.
- Create evaluation frameworks measuring privacy risk and retained data utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
- Ensure systems remain robust to new data sources, schema drift, unusual formats, and hidden sensitive fields.
- Collaborate with engineering, research, operations, and customers to translate privacy requirements into practical technical policies and safeguards.
Required profile
- 2+ years building production data or ML systems in Python.
- Hands-on experience with information extraction, named-entity recognition, classification, or related methods for detecting sensitive content.
- Strong experimental instincts with ability to compare approaches across recall, precision, latency, cost, and downstream utility.
- Experience designing systems that handle schema drift, unusual formats, and edge cases.
- End-to-end experience building data processing pipelines without a fully prescribed roadmap.
Required skills
- Python programming.
- Redaction, masking, pseudonymization, anonymization, and synthetic data generation techniques.
- Familiarity with privacy-enhancing technologies such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption.
- Knowledge of low-latency or high-throughput ML inference and data-processing systems (optional but beneficial).
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in the United States.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 day ago
Expires 1 month from now
4 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
clera
San Francisco