Production ownership
Not just development — production issues, CRs, recovery, and LOB support.
Available for the next hard problem
I’m Nikhil — a Senior Data Engineer and Retail Data Platform Lead simplifying complex production systems, owning the messy issues, and leaving teams with reusable ways to move faster.
Ownership
Large, complex Retail segmentReuse
20+ source pipeline frameworkRecovery
About a week of batch runsLeverage
Two-line prompt workflow01 / why I stand out
The portfolio is about the operating judgment around the pipeline: how complexity gets reduced, how evidence gets built, and how engineering work becomes easier to repeat.
Not just development — production issues, CRs, recovery, and LOB support.
Reusable pipelines and simplification of complex Retail systems.
Controls, business requirements, lineage, and stakeholder interaction.
AI-powered automation and continuous improvement.
01
Big Data02
AWS Data Engineering03
Data Platform Engineering04
Databricks / Lakehouse02 / selected work
Six pieces of work that show how I approach platform complexity: establish evidence, take the ownership, build the repeatable shape, and keep the business in the picture.
Problem: A complex Retail area required broader ecosystem understanding, technical ownership, and less onshore dependency.
What I did: Selected as Retail Lead; simplified workflows, handled business and LOB issues, supported production, shared knowledge, and helped the team operate more independently.
Technology: AWS · PySpark · Spark · SQL · Data platforms
Outcome: Reduced onshore dependency through clearer ownership and simpler ways of working.
Problem: Pipeline health needed to be visible beyond job completion.
What I did: Implemented audit tables with start and end times, source, target, and reject counts; built runtime and completion KPIs and source-based controls to surface sudden drops and anomalies.
Technology: SQL · Data quality controls · Production metrics
Outcome: Better evidence for operational decisions and business issue investigation.
Problem: Scripts required manual path, configuration, migration, and execution changes before they could run in the lab.
What I did: Built an AI agent that understands the script, adapts it for the lab, handles required changes, migrates it, and starts the workflow from a simple two-line prompt.
Technology: AI-powered engineering automation
Outcome: Reduced manual intervention and made production-like testing easier.
Problem: Multiple source integrations created opportunities for duplicate development and inconsistent patterns.
What I did: Built reusable pipelines that emphasize standardization, scalability, and easier onboarding of new sources.
Technology: PySpark · Apache Spark · SQL · Parquet · ORC · Avro
Outcome: A repeatable framework supporting 20+ sources.
Problem: Business stakeholders needed to understand how an opportunity was activated or matched, while production needed recovery after a purge affected about a week of batch runs.
What I did: Innovated lineage and traceability feature tables; investigated HUB issues, unexpected behavior, activation issues, and root causes; worked with BSA and business stakeholders on resolution and recovery.
Technology: Data lineage · Production recovery · Business controls
Outcome: Clearer traceability and stabilized production operations.
Problem: Source-level knowledge silos made it harder for the team to understand the broader domain.
What I did: Initiated a weekly Quick Espresso Call where each team member owned a source, traced its flow, and presented the findings back to the team.
Technology: Source tracing · Data flow understanding · Knowledge sharing
Outcome: Better team independence and more distributed domain knowledge.
03 / profile
A data engineer shaped by production realities: the issue queue, the incident bridge, the source nobody documented, and the business question behind the ticket.
Working thesis
“The platform gets better when the next person can see why.”
Production and business issues need a clear owner, a useful signal, and a path to evidence.
A framework earns its keep when it removes the same decision from tomorrow’s work.
Controls and lineage make technical work explainable to the people relying on it.
AI-powered workflows should create leverage without moving judgment out of the loop.
04 / experience
01
Jun 2025–Present
Client PNC
As Retail Lead, owns a large and complex Retail segment across production reliability, LOB/HUB issue resolution, reusable engineering patterns, and reducing onshore dependency.
Retail Lead · Senior Data Engineer
02
Jan 2024–Jun 2025
Client ixigo
Designed and supported PySpark and Spark ETL workflows across AWS S3, HDFS, EMR, EC2, Hive, Sqoop, and Kafka, with responsibility for cluster monitoring, job orchestration, and streaming data delivery to S3.
Senior Data Engineer
03
Sep 2023–Jan 2024
Client Cisco
Built Spark data-cleansing and ETL workflows, automated incremental Sqoop jobs, and applied partitioning, shuffling, and serialization practices for dependable enterprise data movement.
Data Engineer
04
Aug 2022–Sep 2023
Client Rupeek
Developed distributed Spark applications and end-to-end pipelines using Python, DataFrames, RDDs, Spark SQL, AWS services, YARN monitoring, and Step Functions orchestration.
Data Engineer
05 / technical profile
A production-first toolkit, with Databricks and Lakehouse engineering as the current direction.
Strong programming and problem-solving foundation. Practiced competitive coding on HackerRank · GeeksforGeeks · LeetCode.
Retail Lead, LOB/HUB issue ownership, production recovery, CR deployment ownership, knowledge sharing, and business requirement implementation.
06 / recognition
07 / contact
For hiring conversations, platform problems, or a thoughtful exchange about data engineering.
The contact form will be enabled when the PHP mail handler is added.