Data Engineering
Large-scale ETL and streaming pipelines on Spark, Kafka, and the cloud, plus Teradata & Snowflake warehousing with query tuning and orchestration at production scale.
Data, AI & Software Engineer · Portfolio 2026
Data engineer building large-scale pipelines, warehousing, and cloud infrastructure, and a software & AI engineer bringing LLMs, RAG, and agentic systems to production on top of them. M.S. Computer Science, NYU 2026.
I build the data foundations that production systems run on, large-scale ETL and streaming pipelines, warehousing, and cloud infrastructure, and layer applied AI on top, from LLMs and RAG to multi-agent workflows. I work across the full lifecycle: data preparation and retrieval, model integration, and rigorous evaluation, turning ideas into reliable, real-world software.
Large-scale ETL and streaming pipelines on Spark, Kafka, and the cloud, plus Teradata & Snowflake warehousing with query tuning and orchestration at production scale.
RAG over vector databases with automated retrieval evaluation, LLM fine-tuning and Domain-Adaptive Pre-Training, prompt engineering, and multi-modal generation.
Multi-agent orchestration with LangChain, LangGraph, and LangSmith, with tool-using agents (Claude, Hugging Face, Gemini) for multi-step automation.
Automated validation gates, data quality frameworks, PII protection, and audit trails, aligned to NIST AI risk frameworks and ISO 42001.
CEARTscore · New York, NY
Block Convey · New York, NY
New York University · New York, NY
Accenture Solutions Pvt. Ltd. · Chennai, India
From large-scale data pipelines and cloud infrastructure through to the generative AI and agents that run on top of them.
An agentic pipeline that runs deep web research through Tavily with Perplexity fallback, generates a narrated analysis script via the Claude API, and auto-triggers ElevenLabs TTS to produce finished podcasts delivered to client teams. End to end with no human in the loop. The hard part was defining what “good” means for output nobody can grade.
An MCP-based system that routes AI-proposed code changes through simulated domain-expert personas, tiering scrutiny by risk and catching security, compliance, and team-standard violations before a PR is ever opened. Validated through 70+ B2B customer discovery interviews.
A policy-controlled treasury for autonomous AI agents, with deterministic spending rules and on-chain settlement backed by verifiable audit trails.
Large-scale ETL pipelines feeding Teradata and Snowflake warehouses, built with Ab Initio, Informatica PowerCenter, and Control-M orchestration, with query tuning and stored-procedure design to keep production data processing fast and reliable.
New York University · United States
GPA 3.5 / 4Anna University · Chennai, India
8.2 / 10I’m open to roles and collaborations in data engineering — as well as AI and software engineering. The fastest way to reach me is email.
ap9013@nyu.edu