Skip to main content
Hemanth.
Back to Home

Engineering Experience

Architecting distributed systems, enterprise AI pipelines, and high-throughput backend infrastructure. Click anywhere on any card to explore its architecture case study, trade-offs, and production post-mortems.

AI Python Engineer

Deep DiveCURRENT
LeapGen AI•Ashburn, VA
May 2025 - Present
PythonLangChainLangGraphFastAPIAzure OpenAIGPT-4oDockerTerraformGitHub ActionsMCPRAG

Key Impact & Deliverables

Retrieval-Augmented Generation (RAG) Pipelines

Built Retrieval-Augmented Generation pipelines using LangChain, vector databases, and document embedding/ingestion workflows, improving knowledge-retrieval accuracy and reducing irrelevant chatbot responses.

LangGraph & MCP Workflow Automation

Orchestrated AI workflow automation using LangGraph and Model Context Protocol (MCP), driving document analysis, multi-step reasoning, and task execution across business processes.

Production FastAPI & Azure OpenAI Inference Services

Built production FastAPI services integrating OpenAI GPT-4o and Azure OpenAI into enterprise AI workflows, optimizing prompt orchestration, context management, and API performance for reliable LLM inference.

AI Observability & Performance Tracing

Improved AI application observability by implementing structured logging, tracing, and performance analysis.

Azure Infrastructure & CI/CD Delivery

Deployed and maintained Azure infrastructure using Docker, Terraform, and GitHub Actions, enabling CI/CD delivery.

Seismic & DataCoffee Platform Engineering

Delivered features across Seismic (healthcare platform - billing/subscription management, clinical workflows, system integrations) and DataCoffee (AI data-pipeline builder), owning implementation from backend APIs to production deployment.

Software Engineer

Deep Dive
SmartRevIQ•Delaware, DE
Jun 2023 - Apr 2025
ReactTypeScriptAWS DynamoDBLangGraphAWS S3REST APISSO / RBACTailwind CSSReact Hook Form

Key Impact & Deliverables

High-Throughput Asynchronous Data Pipelines & REST APIs

Built scalable asynchronous data pipelines and REST APIs on AWS DynamoDB with client-side validation, processing thousands of enterprise pricing transactions at sub-3-second query performance for analytics workflows.

Dynamic Schema-Driven Form Engine & Analytics Dashboards

Designed a dynamic form-rendering engine with nested validation and reusable schemas, powering new pricing UI flows across the platform and Built Dashboards with drill-down filtering, surfacing pricing analytics to identify margin leakage.

LangGraph AI Pricing Workflow Automation

Automated enterprise pricing workflows by deploying AI agents with LangGraph, reducing operational overhead.

Enterprise SSO & Role-Based Access Control

Implemented SSO authentication and Role-Based Access Control (RBAC), securing multi-tenant access to sensitive datasets.

Software Engineer

Deep Dive
Speeler Technologies•Hyderabad, India
Apr 2021 – May 2023
ReactGraphQLAWS AppSyncDynamoDBAWS LambdaSQSAWS ECS FargateTensorFlowDockerWebSockets

Key Impact & Deliverables

GraphQL Architecture & Resolver Optimization

Architected and led the redesign of GraphQL backend services for a high-traffic e-commerce platform, implementing DataLoader-style resolver batching - AWS AppSync response caching to cut P95 latency by 30%.

Multi-Tenant DynamoDB Single-Table Architecture

Designed a multi-tenant DynamoDB single-table architecture with sparse GSIs and composite sort keys, eliminating hot-partition bottlenecks and sustaining sub-10ms reads under high-cardinality access patterns, ensuring idempotency.

Event-Driven Image Pipeline & Cost Optimization

Engineered a fault-tolerant, event-driven image-processing pipeline on AWS Lambda and SQS with dead-letter queue handling, cutting infrastructure costs by 70% while processing 2M+ monthly events at 99.99% data durability.

Real-Time Collaborative E-Commerce Platform

Built a real-time collaborative Business Automation & Data Processing E-Commerce Platform with React and WebSockets for concurrent users, reducing perceived latency by 50ms.

Zero-Downtime ECS Fargate Migrations & CI/CD

Migrated EC2 workloads to AWS ECS Fargate with Blue/Green deployments and container-level health checks wired into CI/CD, achieving zero-downtime releases across a 50K+ user production system.

TensorFlow Demand-Forecasting Model

Built a TensorFlow-based demand-forecasting model for e-commerce inventory planning, visualizing forecast trends and accuracy with Matplotlib for stakeholder reporting.

System Architecture Case Study

High-Scale GraphQL & Event-Driven E-Commerce Infrastructure

Business Context & Problem

  • •High-traffic e-commerce platform (50k+ daily requests) with a legacy backend buckling under N+1 query patterns and database hot partitions as tenant count expanded.

System Architecture & Workflow

  • •GraphQL backend featuring DataLoader-style resolver batching and AWS AppSync response caching with tuned TTL policies.
  • •Multi-tenant DynamoDB single-table schema — composite sort keys, sparse GSIs, and write-sharded partition keys.
  • •Event-driven image pipeline utilizing S3, AWS Lambda, SQS Dead Letter Queues (DLQ), and idempotency keys.
  • •React frontend with optimistic UI updates and WebSocket conflict resolution; AWS ECS Fargate container orchestration with Blue/Green deployments.

Why This Architecture Worked

  • ✓Single-table DynamoDB design avoids join overhead and hot-partition issues that occur when running 200+ tenants on shared relational infrastructure.
  • ✓DataLoader batching and AppSync caching directly solved N+1 resolver performance bottlenecks — fixing the true root cause of P95 latency spikes.

Alternatives Considered & Rejected

  • ↳A relational database schema for high-volume access — rejected after load tests proved single-table DynamoDB sustained identical access patterns without hot partitions.
  • ↳Synchronous image processing — rejected as volume scaled; replaced with an asynchronous Lambda + SQS queue pipeline.

Deep Dive: Hard Engineering Realities

Biggest Technical Challenge

Identifying 12 hidden resolver chains exhibiting N+1 patterns — uncovered by profiling production GraphQL query execution plans rather than simple code inspection.

Biggest Production Issue

Synchronous image uploads occasionally backed up the primary web request thread; resolved via Lambda+SQS with DLQs, though initial alerting gaps required refining DLQ monitoring.

Biggest Mistake & Retrospective

Deployed optimistic UI cart updates before conflict-resolution logic was fully hardened — leading to temporary state divergence on concurrent updates; resolved by adding exponential backoff retries and server state reconciliation.

Architectural Trade-offs Accepted:
  • ⚖Single-table DynamoDB provides extreme speed and low cost for pre-planned access patterns, but ad-hoc reporting queries require new GSIs or data migration.
  • ⚖Per-resolver caching required more configuration than full-response caching, but eliminated the risk of serving stale or user-specific shopping cart data.
Scaling Strategy
  • •Sustained sub-10ms p99 reads across 200+ tenants and 8M+ monthly records on single-table DynamoDB.
  • •Processed 2M+ monthly events through the asynchronous image pipeline at 99.99% data durability.
  • •Achieved zero-downtime Blue/Green deployments across 6 production services, replacing 45 minutes of weekly maintenance windows.
Security & Isolation
  • •Idempotency key enforcement on asynchronous event pipelines to prevent duplicate processing or replay attacks.
  • •Partition-level tenant isolation built into DynamoDB primary key structures to guarantee zero cross-tenant data leakage.
Performance Benchmarks
  • •Reduced P95 API latency from 820ms to 570ms (~30% reduction) via resolver batching and AppSync caching.
  • •Achieved sub-10ms p99 DynamoDB read latency and improved perceived cart UI responsiveness by 40%.

Core Engineering Lessons

  • 💡Profiling actual production execution plans is far more effective than guessing bottlenecks from code review.
  • 💡Optimistic UI state conflict resolution is core functionality, not optional polish — it must ship alongside the initial feature.

Common Follow-Up Questions

  • Walk me through what a sparse GSI actually buys you over a normal GSI.
  • How did you pick TTL values for the AppSync cache without serving stale cart or pricing data?
  • What would you do differently on the optimistic-UI rollout, in hindsight?