By clicking the “Apply” button, I understand that my employment application process with Takeda will commence and that the information I provide in my application will be processed in line with Takeda’s Privacy Notice and Terms of Use . I further attest that all information I submit in my employment application is true to the best of my knowledge. Job Description Data Engineering Professional II Digital, Data & Technology (DD&T) - R&D MLOPs About the Role We are seeking an MLOps Engineer to operationalize machine learning and generative AI across our R&D and enterprise data ecosystem. You will build and maintain the platforms, pipelines, and controls that move models from notebook experiments into validated, production-grade, GxP-compliant services: supporting use cases that span clinical development, regulatory operations, pharmacovigilance, real-world evidence, and translational/biomarker research. This is a hands-on engineering role at the intersection of data engineering, ML lifecycle automation, and regulated-systems discipline. You will work primarily in Databricks and AWS, partnering with data scientists, platform/cloud engineering, quality, and regulatory teams to ship models that are reproducible, monitored, auditable, and trustworthy. Key Responsibilities ML Lifecycle & Pipeline Automation Design, build, and operate end-to-end ML pipelines (data ingestion → feature engineering → training → validation → deployment → monitoring) using Databricks (Delta Lake, MLflow, Unity Catalog, Feature Store, Workflows/Jobs) and AWS services. Implement CI/CD for ML and data assets (e.g., GitHub Actions, GitLab CI, or Jenkins), including automated testing, environment promotion (dev → test → prod), and reproducible builds. Stand up and maintain model registries, model versioning, and artifact lineage so every deployed model is traceable to its data, code, and configuration. Cloud & Platform Engineering (AWS) Build and manage ML infrastructure on AWS — e.g., SageMaker, Bedrock, S3, Lambda, ECS/EKS, Step Functions, ECR, IAM, CloudWatch — using Infrastructure as Code (Terraform or CloudFormation/CDK). Integrate Databricks with AWS securely (Unity Catalog governance, cross-account access, VPC/networking, KMS encryption, secrets management). Optimize compute and cost (cluster policies, autoscaling, spot strategy, job orchestration) without compromising performance or compliance. Production Monitoring & Reliability Implement model and data monitoring: drift detection, data-quality checks, performance/SLA tracking, and automated alerting/retraining triggers. Establish observability and incident-response practices for ML services; participate in on-call/runbook ownership as needed. Maintain feature stores and data contracts to ensure consistency between training and serving. Regulated-Environment & Compliance Engineering Build ML systems that meet GxP expectations and support Computer System Validation (CSV) / Computer Software Assurance (CSA), GAMP 5, 21 CFR Part 11, and data-integrity (ALCOA+) requirements. Implement audit trails, electronic records/signatures controls, access controls, and change-management workflows suitable for validated environments. Handle PII/PHI and sensitive R&D data in line with HIPAA, GDPR, and internal privacy/data-governance policies (de-identification, anonymization, role-based access). Author and maintain technical documentation, validation deliverables, and SOP-aligned procedures; partner with Quality/QA and Regulatory on audits and inspections. Collaboration & Enablement Work under the guidance of Director, Solution Engineering/Solution Architect to produce artifacts and deliverables that adhere to best practices at Takeda. Partner with data scientists to productionize models (including LLM/GenAI and RAG applications) and to translate research code into robust, maintainable services. Contribute reusable templates, accelerators, and self-service tooling that raise the engineering bar across teams. Promote MLOps best practices, mentor peers, and document standards. Required Qualifications Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience). 4+ years of hands-on experience in MLOps, ML engineering, data engineering, or DevOps for data/ML systems. Strong Databricks experience: Delta Lake, MLflow, Unity Catalog, Jobs/Workflows, and Spark (PySpark). Strong AWS experience across compute, storage, and IAM, plus at least one ML service (SageMaker and/or Bedrock). Proficiency in Python for production code (packaging, testing, typing), plus solid SQL. Experience building CI/CD pipelines and using Git-based workflows. Working knowledge of containerization (Docker) and orchestration (Kubernetes/EKS or ECS). Experience with Infrastructure as Code (Terraform, CloudFormation, or CDK). Understanding of ML lifecycle concepts: experiment tracking, model registry, feature stores, and model monitoring/drift. Preferred / Pharma-Specific Qualifications Experience delivering software or ML in a GxP / regulated (FDA, EMA) life-sciences environment; familiarity with CSV/CSA, GAMP 5, 21 CFR Part 11, ALCOA+. Exposure to pharma/biotech data domains: clinical trial data (CDISC/SDTM/ADaM), regulatory submissions, pharmacovigilance/safety, real-world data (RWD/RWE), or omics/biomarker datasets. Experience operationalizing LLM/GenAI workloads (e.g., AWS Bedrock), including RAG, prompt/version management, evaluation, and guardrails. Familiarity with handling PHI/PII under HIPAA/GDPR and with data-governance tooling. Streaming/event-driven data (Kafka/Kinesis), data-observability tooling, and feature-store frameworks. Relevant certifications: AWS (ML Specialty, Solutions Architect, or DevOps Engineer) and/or Databricks (Data Engineer, ML Engineer). What Success Looks Like (First 12 Months) Production ML/GenAI pipelines run reproducibly with full lineage, monitoring, and automated promotion across validated environments. Deployment lead time and manual handoffs are measurably reduced through reusable templates and CI/CD. Models in production are monitored for drift and quality, with documented retraining and rollback procedures. Engineering artifacts meet inspection-readiness standards and pass internal QA review. Locations IND - Bengaluru Worker Type Employee Worker Sub-Type Regular Time Type Full time
Data Engineering Professional II
Takeda
JP
🎯 Mid-level📄 Permanent🏠 On-site🧭 Ml-ai🏢 Healthcare🗣️ English
Required skills
databricksawsmlflowdelta lakefeature storegithub actionsgitlab cijenkinssagemaker
Ready to interview?
Upload your CV on tailk.me and let the AI run the interview for you.
Build your AI headhunter and interviewRelated listings
GCM Data Dissemination Lead
Takeda
IND - Bengaluru
Data Engineering Professional I
Takeda
IND - Bengaluru
Business Analyst
Takeda
IND - Bengaluru
Senior Manager, Decision Science
Takeda
IND - Bengaluru
Global Subchapter Lead – Service Platform Engineering
Takeda
IND - Bengaluru · Remote
Associate Director, Computational Biology
Takeda
IND - Bengaluru - Research and Development
tailk