Senior Data Scientist

About Providence

Providence, one of the US’s largest not-for-profit healthcare systems, is committed to high quality, compassionate healthcare for all. Driven by the belief that health is a human right and the vision, ‘Health for a better world’, Providence and its 121,000 caregivers strive to provide everyone access to affordable quality care and services.

Providence has a network of 51 hospitals, 1,000+ care clinics, senior services, supportive housing, and other health and educational services in the US.

Providence India is bringing to fruition the transformational shift of the healthcare ecosystem to Health 2.0. The India center will have focused efforts around healthcare technology and innovation, and play a vital role in driving digital transformation of health systems for improved patient outcomes and experiences, caregiver efficiency, and running the business of Providence at scale.


Why Us?

  • Best In-class Benefits
  • Inclusive Leadership
  • Reimagining Healthcare
  • Competitive Pay
  • Supportive Reporting Relation

Job Description: Senior Data Scientist (Statistics, ML & Generative AI)

Location: Hyderabad, Telangana, India / Hybrid

Role Type: Full-Time

Position Overview

We are looking for a highly analytical and technical Senior Data Scientist with a strong mathematical foundation in statistics, probability, and classical machine learning, paired with cutting-edge experience in Large Language Models (LLMs) and Generative AI.

In this role, you will bridge the gap between rigorous statistical modeling and state-of-the-art AI technologies. You will design, develop, and deploy end-to-end data science solutions, RAG (Retrieval-Augmented Generation) architectures, and predictive models that turn complex data into strategic business value.

Key Responsibilities

  • Core Statistical & ML Modeling:

    • Apply advanced probability theory, hypothesis testing, Bayesian methods, and statistical inference to validate data distributions, perform predictive analytics, and design robust experiments (A/B testing).
    • Build, fine-tune, and evaluate classical machine learning algorithms (regression, classification, clustering, time-series forecasting, ensemble methods) for production systems.
  • Generative AI & LLM Engineering:

    • Design and implement Generative AI architectures, including Retrieval-Augmented Generation (RAG) pipelines, multi-agent frameworks, and vector search systems.
    • Fine-tune, evaluate, and prompt-engineer open-source and proprietary Large Language Models (LLMs) for specific domain applications.
    • Implement guardrails, evaluation frameworks (e.g., RAGAS, DeepEval), and toxicity/bias detection for LLM applications.
  • Data Engineering & System Architecture:

    • Collaborate with engineering teams to deploy ML/LLM models into production microservices via REST APIs, gRPC, or async worker queues.
    • Optimize model inference latency, throughput, and GPU/memory consumption.
    • Design data pipelines using Python, SQL, and modern data orchestration toolkits.
  • Thought Leadership & Collaboration:

    • Partner with product and business stakeholders to translate complex business problems into scalable machine learning and AI frameworks.
    • Stay at the forefront of emerging research in machine learning, deep learning, and generative AI.

Required Qualifications & Technical Skills

  • Education: Bachelor's, Master's, or Ph.D. in Computer Science, Statistics, Applied Mathematics, Data Science, Quantitative Finance, or a related quantitative field.
  • Core Mathematics & Statistics:

    • Strong expertise in probability distributions, stochastic processes, decision theory, regression analysis, and variance reduction techniques.
  • Classical ML & Deep Learning:

    • Proficiency with ML frameworks (scikit-learn, XGBoost, LightGBM, PyTorch, or TensorFlow).
    • Deep understanding of loss functions, optimization algorithms, evaluation metrics (ROC-AUC, RMSE, Precision/Recall, BLEU, ROUGE), and cross-validation techniques.
  • Generative AI & LLM Ecosystem:

    • Practical experience with LLM frameworks like LangChain, LangGraph, LlamaIndex, or DSPy.
    • Hands-on experience with Vector Databases (ChromaDB, Pinecone, pgvector, Qdrant, or FAISS).
    • Knowledge of quantization techniques (LoRA, QLoRA), embedding models, and multi-agent coordination frameworks (e.g., Model Context Protocol / MCP servers).
  • Programming & Tools:

    • High proficiency in Python (Pandas, NumPy, SciPy, PyTorch) and complex SQL.
    • Familiarity with containerization (Docker) and modern development workflows (Git, CI/CD).

Preferred Qualifications

  • Experience with time-series forecasting, quantitative modeling, or risk prediction algorithms.
  • Knowledge of MLOps practices, model monitoring, and pipeline orchestration tools.
  • Familiarity with cloud platforms (AWS, GCP, or Azure).

Providence’s vision to create ‘Health for a Better World’ aids us to provide a fair and equitable workplace for all in our employment, whether temporary, part-time or full time, and to promote individuality and diversity of thought and background, and acknowledge its role in the organization’s success. This makes us committed towards equal employment opportunities, regardless of race, religion or belief, color, ancestry, disability, marital status, gender, sexual orientation, age, nationality, ethnic origin, pregnancy, or related needs, mental or sensory disability, HIV Status, or any other category protected by applicable law. In furtherance to our mission in building a more inclusive and equitable environment, we shall, from time to time, undertake programs to assist, uplift and empower underrepresented groups including but not limited to Women, PWD (Persons with Disabilities), LGTBQ+ (Lesbian, Gay, Transgender, Bisexual or Queer), Veterans and others. We strive to address all forms of discrimination or harassment and provide a safe and confidential process to report any misconduct.

Contact our Integrity hotline also, read our Code of Conduct.