Railway Predictive Maintenance Platform
-
Led the architecture and delivery of a large-scale predictive maintenance data platform for a major railway operator, processing high-volume telemetry and time-series data using Apache Spark deployed on Kubernetes (AWS). Acted as technical lead and solution architect, working closely with product owners, data scientists, and full-stack teams to deliver a production-grade, ML-ready data platform.
Key responsibilities and achievements
-
Designed and implemented end-to-end data architecture covering ingestion, batch and streaming processing, feature engineering, and analytics consumption.
-
Built, optimized, and maintained Apache Spark (Scala, Python) pipelines, focusing on performance, reliability, and scalability.
-
Deployed and operated Spark workloads on Kubernetes (AWS), leveraging CI/CD (Gitlab) and infrastructure-as-code practices.
- Led functional and technical monitoring of the data platform, defined data reliability and production standards, and served as an escalation point for incidents, including:
- Job health and SLA tracking,
- Data quality controls and anomaly detection,
- Dataset availability and freshness guarantees,
- Distributed tracing and observability across Spark pipelines using Datadog.
-
Partnered closely with data scientists to productionize predictive models and ensure scalable training and inference workflows.
-
Worked daily with product owners, backend/frontend engineers, translating operational and business requirements into robust data solutions.
- Conducted architecture reviews, technical workshops, and design trade-off discussions with stakeholders.