Hi, I'm Tharun. I build ML systems by day and job-hunting robots by night.

The day job is data science at AT&T. The night project is an agent that fills out job applications on its own — and knows exactly when to wake me up.

AT&T logoAT&T Stanford logoStanford Medicine LTIMindtree logoLTIMindtree SFSU logoSFSU

What brings you here?

or just scroll — I'll narrate as you go

Tharun Kumar Reddy Byreddy standing at a mountain lake off duty · zero bars of signal

tharun> status --now

Right now

A portfolio is a snapshot; this is the live feed.

Building

The job application agent below — currently wiring up the Gmail confirmation watcher and hardening the checkpoint flow.

Learning

Deep in an 8-week agentic AI sprint: LangGraph, CrewAI, and LlamaIndex courses, plus Google Cloud Skills Boost labs on Vertex AI and Terraform.

Reading

Anthropic's engineering posts on building effective agents. Strong opinions on when not to use one.

tharun> projects --order-by shipped_last

Things I've built

Nine public repos below, plus the work-in-progress I can't stop tinkering with.

SQL Analytics Agent

Ask a question in English, get an answer backed by SQL. An Anthropic SDK tool-use loop that writes queries, runs them, reads results, and keeps going until the question is actually answered.

Anthropic SDK · SQLite · Streamlit view source →

RAG Document Q&A

Document question-answering with retrieval done right — LangChain orchestration, ChromaDB vector store, answers grounded in retrieved context instead of vibes.

LangChain · ChromaDB · embeddings view source →

TransplantMatch

Predicting one-year post-transplant outcomes — death, graft failure, re-transplantation — with Cox PH, logistic regression, and random forests. Medicine keeps the statistics honest.

Cox PH · random forest · survival analysis view source →

Hospital Readmission Predictor

30-day readmission risk with SHAP explainability, so a clinician can see why a patient scores high — served as a Streamlit app.

scikit-learn · SHAP · Streamlit view source →

Clinical Survival Analysis

Cox proportional hazards on clinical data with an interactive dashboard for hazard ratios, survival curves, and covariate effects.

Cox PH · Streamlit · statistics view source →

Fraud Detection

Classification on heavily imbalanced transaction data — the classic problem where accuracy lies to you and precision-recall tradeoffs decide if the model is useful.

XGBoost · imbalanced learning view source →

Time-Series Forecasting

Hybrid LSTM + statistical forecasting for operational metrics, with anomaly detection and seasonality decomposition.

LSTM · ARIMA · Prophet view source →

ML Pipeline Automation

Feature engineering, model comparison, and experiment tracking wired into one reproducible pipeline. One command, not a folder of notebooks.

Python · experiment tracking view source →

SQL Analytics Case Study

Advanced SQL on NYC Taxi data, written the way analysis actually happens: business question, query, finding, recommendation.

SQL · window functions · Python viz view source →

tharun> experience --reverse-chronological

Where I've worked

Telecom scale, academic medicine, enterprise consulting. Three very different data cultures; the constant is shipping work people act on.

AT&T

AT&T

SEP 2024 — PRESENT

San Francisco, CA

Data Scientist

  • Churn risk models (XGBoost, LightGBM) on 150M+ daily clickstream events — 0.91 AUC, 85% recall, and a 14% drop in customer churn.
  • Collaborative filtering recommendations across web, mobile, and email. CTR up 24%, offer conversion up 12-15%.
  • 10+ A/B experiments where statistical significance turned into deployment decisions, wired into a Snowflake feedback loop.
  • Power BI dashboards executives actually open — impressions, clicks, conversions, revenue attribution.
Stanford

Stanford Medicine

SEP 2023 — AUG 2024

Redwood City, CA

Research Data Scientist

  • A multi-agent RAG co-pilot (LangChain, GPT-4) that reads oncology literature so researchers don't have to — review time down ~70% across 100+ live clinical queries.
  • Hybrid semantic search, BM25 plus embeddings in FAISS: 85%+ relevance precision, 4.5/5 clinician rating on summaries.
  • Survival models on heart transplant registries (DHS, SRTR) — Cox PH, mixed effects, and a random forest at 92% accuracy.
  • Event-driven GCP pipelines keeping the evidence base fresh in under 24 hours.
LTIMindtree

LTIMindtree

SEP 2021 — JUL 2022

Hyderabad, India

Data Analyst

  • SQL/Python ETL on 1M+ daily transaction records, holding 99%+ accuracy.
  • 10+ Power BI dashboards that retired a stack of manual Excel reports.
  • Reporting automation that gave the team back most of a day every week.

tharun> skills --group-by domain

Toolkit

Colored pills are what I reach for most right now.

GenAI & agents

  • LangGraph
  • LangChain
  • RAG
  • Anthropic SDK
  • Vector embeddings
  • Semantic search
  • Prompt engineering
  • Playwright
  • Multi-agent frameworks

ML & statistics

  • Python
  • scikit-learn
  • XGBoost
  • LSTM / ARIMA / Prophet
  • A/B testing
  • R

Data, cloud & BI

  • SQL
  • PostgreSQL
  • BigQuery
  • Snowflake
  • Apache Spark
  • Airflow
  • AWS (Secrets Manager)
  • GCP
  • Docker
  • Streamlit
  • Power BI
  • Tableau

tharun> education

Education

San Francisco State University

M.S. Statistical Data Science — San Francisco State University, 2024

Multivariate Statistics · Regression Analysis · Applied Machine Learning · Statistical Computing · Time Series Analysis

Also tutored statistics through SFSU's SSS-TRiO program — the best test of whether you understand regression is explaining it to someone seeing it for the first time.

3.77GPA / 4.00

tharun> contact --channel any

Say hello

Open to Data Scientist, ML Engineer, and applied AI roles. If you're building something where production ML meets agents, I want to hear about it.

© 2026 Tharun Kumar Reddy Byreddy · San Francisco, CA · last updated July 2026 no template was harmed — view source if you don't believe me