Senior Data Engineer
Hybrid
Ascendion
Enterprise
Product & Service
B2B
₹ 25-27 Lacs PA
Pre-seed
Information Technology
Pune, Maharashtra, India
Post Status: Active
Permanent
12 applications
Experience: 8-15 Years
Skills
Apache Spark
CI/CD
SQL
BigQuery
Apache Kafka
Docker
Machine Learning
Kubernetes
Python
PySpark
MLOps
US Health Care
Google Cloud Platform (GCP)
Vertex AI
Posted 14 days ago

About the job

US Healthcare | Data Engineering | Machine Learning Engineering

We are looking for an experienced Senior Data Engineer / Data Engineer II with strong expertise in Google Cloud Platform (GCP) and Machine Learning Engineering to build and manage scalable data platforms and ML pipelines for enterprise healthcare applications. The ideal candidate should have hands-on experience in designing cloud-native data solutions, developing feature engineering pipelines, supporting MLOps workflows, and enabling AI/ML models in production.

Key Responsibilities

  • Design, develop, and maintain scalable data pipelines on Google Cloud Platform (GCP).

  • Build batch and streaming data pipelines using Apache Spark (PySpark), Dataflow, and Kafka.

  • Develop and optimize data models and ETL/ELT workflows for large-scale datasets.

  • Work extensively with BigQuery for data warehousing, analytics, and performance optimization.

  • Design and maintain ML data pipelines supporting model training and inference.

  • Build and manage Feature Stores for machine learning applications.

  • Deploy and manage ML workflows using Vertex AI and MLOps best practices.

  • Collaborate with Data Scientists, ML Engineers, and cross-functional teams to productionize ML models.

  • Implement CI/CD practices for data and ML pipelines.

  • Ensure data quality, governance, monitoring, and reliability across cloud environments.

  • Optimize pipeline performance, scalability, and cost efficiency.

Mandatory Skills

  • Strong hands-on experience with Google Cloud Platform (GCP)

  • BigQuery

  • Apache Spark (PySpark)

  • Dataflow

  • Kafka

  • Vertex AI

  • Python and/or Java

  • SQL

  • ML Data Pipelines

  • Feature Store

  • MLOps

  • ETL/ELT Pipeline Development

  • Cloud Data Architecture

Preferred Skills

  • Experience in the US Healthcare domain.

  • Understanding of healthcare data standards and large-scale healthcare data platforms.

  • Knowledge of CI/CD pipelines for data and ML workloads.

  • Experience with containerization (Docker/Kubernetes) is an added advantage.

  • Familiarity with modern AI/ML lifecycle management and deployment practices.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.

  • 8+ years of overall IT experience.

  • Strong experience in GCP-based Data Engineering and Machine Learning Engineering.

  • Excellent analytical, problem-solving, and communication skills.

Preferred Candidate Profile

  • Immediate Joiners

  • Candidates serving notice period (preferred joining within 20–30 days)

  • Hands-on experience in building production-grade GCP Data Engineering and ML pipelines.

  • Experience working in Agile development environments.

Interview Process

  • Round 1: Technical Interview (Virtual)

  • Round 2: Technical/Managerial Interview

  • Round 3: Client Discussion (if applicable)