Data Scientist · Current
AMD
Building production machine learning systems from large-scale telemetry, with a focus on reliability, prediction, and practical engineering impact.
01 · Context
Working where machine learning meets hardware.
At AMD, my work sits at the intersection of machine learning, large-scale telemetry, and production engineering. I work with signals generated by GPU and CPU systems and turn them into models that help teams anticipate failures, understand abnormal behavior, and make complex infrastructure more reliable.
The work is deeply collaborative. It involves partnering with data engineering and platform teams, because model quality depends as much on reliable pipelines, monitoring, and deployment as it does on the algorithm itself.
02 · Predictive systems
From telemetry to early warning.
I built predictive models using LightGBM and PySpark across telemetry from more than 18,000 production nodes. The system was designed to identify the likelihood of GPU or CPU failure within a 30-day window and reached an AUROC of 0.88.
A large part of the work involved engineering time-series features from ECC errors, thermal signals, and PCIe telemetry. That feature work improved model performance by six percentage points in AUROC and made the pipeline suitable for distributed processing.
03 · Anomaly detection
Finding patterns that rules miss.
I also worked on anomaly detection for manufacturing telemetry using XGBoost and Isolation Forest. Compared with rule-based methods, the approach improved defect detection by 18% and gave reliability teams a faster path to root-cause analysis.
This reinforced something I value in applied data science: the most useful model is not simply the most accurate one. It is the one that gives teams a clearer and earlier signal they can act on.
04 · Production ML
Making models observable and dependable.
To support production reliability, I established an MLflow model registry and integrated PSI-based drift monitoring into Airflow scoring pipelines. This created a clearer model lifecycle—from experimentation and registration to monitoring and governance.
The goal was to make model behavior visible over time, not treat deployment as the end of the work.
05 · Language models
Using technical language as data.
I fine-tuned language models with LoRA and PEFT on hardware failure logs, reaching 89% classification accuracy and automating ticket tagging across more than 10,000 monthly support cases.
I also designed retrieval-augmented generation pipelines over hardware documentation and failure reports using LangChain, FAISS, and vector search. The system reduced diagnosis time by 32% compared with traditional keyword search.
06 · What stays with me
Systems thinking over isolated models.
This role has shaped how I think about data science. Strong results come from the full system: thoughtful features, scalable data pipelines, careful evaluation, clear monitoring, and close collaboration with the people who use the output.
It is this combination of modeling and engineering that I want to keep exploring in my future work.
07 · The next question
Can predictive systems become useful partners rather than silent alarms?
The next step I want to explore is how forecasting, anomaly detection, retrieval, and explanation can work together. A system should not only flag that something may fail; it should help an engineer understand what changed, what evidence matters, and what action is worth considering.
