Data, AI & Software Engineer · Portfolio 2026

AkshayaaPrabakar

Data engineer building large-scale pipelines, warehousing, and cloud infrastructure, and a software & AI engineer bringing LLMs, RAG, and agentic systems to production on top of them. M.S. Computer Science, NYU 2026.

About

Data systems first — intelligence on top.

I build the data foundations that production systems run on, large-scale ETL and streaming pipelines, warehousing, and cloud infrastructure, and layer applied AI on top, from LLMs and RAG to multi-agent workflows. I work across the full lifecycle: data preparation and retrieval, model integration, and rigorous evaluation, turning ideas into reliable, real-world software.

3+Years building data & AI systems
~300Grad students mentored as Head TA
~40KDocuments processed for LLM training
1Startup co-founded (NYU Leslie Entrepreneurial Institute)

Data Engineering

Large-scale ETL and streaming pipelines on Spark, Kafka, and the cloud, plus Teradata & Snowflake warehousing with query tuning and orchestration at production scale.

Applied GenAI & LLMs

RAG over vector databases with automated retrieval evaluation, LLM fine-tuning and Domain-Adaptive Pre-Training, prompt engineering, and multi-modal generation.

Agentic Systems

Multi-agent orchestration with LangChain, LangGraph, and LangSmith, with tool-using agents (Claude, Hugging Face, Gemini) for multi-step automation.

Trustworthy ML & Governance

Automated validation gates, data quality frameworks, PII protection, and audit trails, aligned to NIST AI risk frameworks and ISO 42001.

Experience

Three years across data, AI & engineering.

  1. Software Engineer (Intern to Full Time)

    CEARTscore · New York, NY

    Jan 2026 – Present
    • Built a RAG system over vector databases with embedding workflows, and multi-agent LLM workflows that automated multi-step document analysis and content generation, cutting manual analysis effort significantly.
    • Designed and automated a multi-source data pipeline on Azure processing ~1,000 documents, with PII reduction, validation, secure storage, and preprocessing for model training.
    • Prepared and processed a corpus of ~40K documents for domain-specific LLM fine-tuning and Domain-Adaptive Pre-Training.
    • Built an automated evaluation harness that runs on each pipeline execution, scoring RAG retrieval against known-good answers and surfacing what passed, what failed, and the retrieved output beside the expected one.
    RAGLLM AgentsAzureFine-TuningEvals
  2. Software Engineering Intern

    Block Convey · New York, NY

    Jun 2025 – Aug 2025
    • Built an end-to-end synthetic data generator with multi-stage PII masking; a core contribution to the company’s patent portfolio.
    • Re-engineered a monolithic pipeline into a parallel, agent-based system on GCP, roughly doubling processing throughput (~50% faster) via distributed workflow handling, validation, and monitoring.
    • Developed a NIST-aligned data validation and quality framework that reduced manual review and reporting time by ~80% through automated pre-validation.
    Synthetic DataMulti-AgentGCPAI Governance
  3. Graduate Teaching Assistant, Big Data (CS-GY-6513)

    New York University · New York, NY

    May 2025 – May 2026
    • Head TA for the course; mentored nearly 300 graduate students in Hadoop, Spark, Dask, Kafka, and NoSQL, held TA hours and clarified queries, evaluated assignments with targeted feedback, and guided 200+ project teams from proposal to final demonstration.
    SparkKafkaHadoopMentorship
  4. Data Engineering Associate

    Accenture Solutions Pvt. Ltd. · Chennai, India

    Oct 2022 – Jul 2024
    • Optimized data transformation processes using Ab Initio and Informatica PowerCenter, improving data accuracy and ETL performance by ~40% through advanced Teradata query tuning, and streamlined dozens of Control-M jobs to enhance ETL workflows and reduce processing errors, ensuring reliable validation, migration, and testing.
    • Designed and refined stored procedures, views, and tables in Teradata and Snowflake, improving query performance to support data-driven decision-making.
    ETLTeradataSnowflakeControl-M
Capabilities

A full-stack data & AI toolkit.

From large-scale data pipelines and cloud infrastructure through to the generative AI and agents that run on top of them.

Data Engineering

  • SQL / PL-SQL
  • Apache Spark
  • PySpark
  • Hadoop
  • Kafka
  • Dask
  • Ab Initio
  • Informatica PowerCenter
  • Control-M
  • Teradata
  • Snowflake

Cloud & DevOps

  • AWS (S3, EC2, Lambda)
  • Azure
  • GCP
  • Docker
  • Kubernetes
  • Jenkins
  • CI/CD
  • Git / GitHub

AI & GenAI

  • Large Language Models
  • RAG
  • LLM Fine-Tuning
  • Domain-Adaptive Pre-Training
  • Prompt Engineering
  • Evaluation Harnesses
  • Multi-Modal Generation
  • Embeddings
  • Vector Databases

Agentic Systems

  • LangChain
  • LangGraph
  • LangSmith
  • Multi-Agent Orchestration
  • Tool-Using Agents
  • Claude
  • Hugging Face
  • Gemini

Machine Learning

  • TensorFlow
  • PyTorch
  • Model Evaluation
  • Synthetic Data Generation

Languages & Governance

  • Python
  • SQL
  • Bash / Unix
  • NIST AI Risk
  • ISO 42001
  • PII Protection
  • SC-900
Selected Work

Things I’ve built.

01
Production · Agentic

Research-to-Podcast Pipeline

An agentic pipeline that runs deep web research through Tavily with Perplexity fallback, generates a narrated analysis script via the Claude API, and auto-triggers ElevenLabs TTS to produce finished podcasts delivered to client teams. End to end with no human in the loop. The hard part was defining what “good” means for output nobody can grade.

Claude APITavilyElevenLabsMulti-AgentProduction
02
Co-founded · NYU Leslie Entrepreneurial Institute

AI Context Engine

An MCP-based system that routes AI-proposed code changes through simulated domain-expert personas, tiering scrutiny by risk and catching security, compliance, and team-standard violations before a PR is ever opened. Validated through 70+ B2B customer discovery interviews.

MCPMulti-AgentLangGraphRisk TieringGovernance
03
MIT Bitcoin Hackathon, Community Prize 🏆

Eclipse

A policy-controlled treasury for autonomous AI agents, with deterministic spending rules and on-chain settlement backed by verifiable audit trails.

Autonomous AgentsOn-Chain SettlementAudit Trails
04
Data Engineering · Production

Enterprise ETL & Warehouse Optimization

Large-scale ETL pipelines feeding Teradata and Snowflake warehouses, built with Ab Initio, Informatica PowerCenter, and Control-M orchestration, with query tuning and stored-procedure design to keep production data processing fast and reliable.

TeradataSnowflakeAb InitioInformaticaControl-M
Aug 2024 – May 2026

M.S. in Computer Science

New York University · United States

GPA 3.5 / 4
Aug 2018 – May 2022

B.E. in Computer Science & Engineering

Anna University · Chennai, India

8.2 / 10
Certification

SC-900, Security, Compliance & Identity

Let’s build

Have an idea worth engineering?

I’m open to roles and collaborations in data engineering — as well as AI and software engineering. The fastest way to reach me is email.

ap9013@nyu.edu