Engineering Experience
Architecting distributed systems, enterprise AI pipelines, and high-throughput backend infrastructure. Click anywhere on any card to explore its architecture case study, trade-offs, and production post-mortems.
AI Python Engineer
Deep DiveCURRENTKey Impact & Deliverables
Built Retrieval-Augmented Generation pipelines using LangChain, vector databases, and document embedding/ingestion workflows, improving knowledge-retrieval accuracy and reducing irrelevant chatbot responses.
Orchestrated AI workflow automation using LangGraph and Model Context Protocol (MCP), driving document analysis, multi-step reasoning, and task execution across business processes.
Built production FastAPI services integrating OpenAI GPT-4o and Azure OpenAI into enterprise AI workflows, optimizing prompt orchestration, context management, and API performance for reliable LLM inference.
Improved AI application observability by implementing structured logging, tracing, and performance analysis.
Deployed and maintained Azure infrastructure using Docker, Terraform, and GitHub Actions, enabling CI/CD delivery.
Delivered features across Seismic (healthcare platform - billing/subscription management, clinical workflows, system integrations) and DataCoffee (AI data-pipeline builder), owning implementation from backend APIs to production deployment.
Software Engineer
Deep DiveKey Impact & Deliverables
Built scalable asynchronous data pipelines and REST APIs on AWS DynamoDB with client-side validation, processing thousands of enterprise pricing transactions at sub-3-second query performance for analytics workflows.
Designed a dynamic form-rendering engine with nested validation and reusable schemas, powering new pricing UI flows across the platform and Built Dashboards with drill-down filtering, surfacing pricing analytics to identify margin leakage.
Automated enterprise pricing workflows by deploying AI agents with LangGraph, reducing operational overhead.
Implemented SSO authentication and Role-Based Access Control (RBAC), securing multi-tenant access to sensitive datasets.
Software Engineer
Deep DiveKey Impact & Deliverables
Architected and led the redesign of GraphQL backend services for a high-traffic e-commerce platform, implementing DataLoader-style resolver batching - AWS AppSync response caching to cut P95 latency by 30%.
Designed a multi-tenant DynamoDB single-table architecture with sparse GSIs and composite sort keys, eliminating hot-partition bottlenecks and sustaining sub-10ms reads under high-cardinality access patterns, ensuring idempotency.
Engineered a fault-tolerant, event-driven image-processing pipeline on AWS Lambda and SQS with dead-letter queue handling, cutting infrastructure costs by 70% while processing 2M+ monthly events at 99.99% data durability.
Built a real-time collaborative Business Automation & Data Processing E-Commerce Platform with React and WebSockets for concurrent users, reducing perceived latency by 50ms.
Migrated EC2 workloads to AWS ECS Fargate with Blue/Green deployments and container-level health checks wired into CI/CD, achieving zero-downtime releases across a 50K+ user production system.
Built a TensorFlow-based demand-forecasting model for e-commerce inventory planning, visualizing forecast trends and accuracy with Matplotlib for stakeholder reporting.
High-Scale GraphQL & Event-Driven E-Commerce Infrastructure
Business Context & Problem
- •High-traffic e-commerce platform (50k+ daily requests) with a legacy backend buckling under N+1 query patterns and database hot partitions as tenant count expanded.
System Architecture & Workflow
- •GraphQL backend featuring DataLoader-style resolver batching and AWS AppSync response caching with tuned TTL policies.
- •Multi-tenant DynamoDB single-table schema — composite sort keys, sparse GSIs, and write-sharded partition keys.
- •Event-driven image pipeline utilizing S3, AWS Lambda, SQS Dead Letter Queues (DLQ), and idempotency keys.
- •React frontend with optimistic UI updates and WebSocket conflict resolution; AWS ECS Fargate container orchestration with Blue/Green deployments.
Why This Architecture Worked
- ✓Single-table DynamoDB design avoids join overhead and hot-partition issues that occur when running 200+ tenants on shared relational infrastructure.
- ✓DataLoader batching and AppSync caching directly solved N+1 resolver performance bottlenecks — fixing the true root cause of P95 latency spikes.
Alternatives Considered & Rejected
- ↳A relational database schema for high-volume access — rejected after load tests proved single-table DynamoDB sustained identical access patterns without hot partitions.
- ↳Synchronous image processing — rejected as volume scaled; replaced with an asynchronous Lambda + SQS queue pipeline.
Deep Dive: Hard Engineering Realities
Identifying 12 hidden resolver chains exhibiting N+1 patterns — uncovered by profiling production GraphQL query execution plans rather than simple code inspection.
Synchronous image uploads occasionally backed up the primary web request thread; resolved via Lambda+SQS with DLQs, though initial alerting gaps required refining DLQ monitoring.
Deployed optimistic UI cart updates before conflict-resolution logic was fully hardened — leading to temporary state divergence on concurrent updates; resolved by adding exponential backoff retries and server state reconciliation.
- ⚖Single-table DynamoDB provides extreme speed and low cost for pre-planned access patterns, but ad-hoc reporting queries require new GSIs or data migration.
- ⚖Per-resolver caching required more configuration than full-response caching, but eliminated the risk of serving stale or user-specific shopping cart data.
Scaling Strategy
- •Sustained sub-10ms p99 reads across 200+ tenants and 8M+ monthly records on single-table DynamoDB.
- •Processed 2M+ monthly events through the asynchronous image pipeline at 99.99% data durability.
- •Achieved zero-downtime Blue/Green deployments across 6 production services, replacing 45 minutes of weekly maintenance windows.
Security & Isolation
- •Idempotency key enforcement on asynchronous event pipelines to prevent duplicate processing or replay attacks.
- •Partition-level tenant isolation built into DynamoDB primary key structures to guarantee zero cross-tenant data leakage.
Performance Benchmarks
- •Reduced P95 API latency from 820ms to 570ms (~30% reduction) via resolver batching and AppSync caching.
- •Achieved sub-10ms p99 DynamoDB read latency and improved perceived cart UI responsiveness by 40%.
Core Engineering Lessons
- 💡Profiling actual production execution plans is far more effective than guessing bottlenecks from code review.
- 💡Optimistic UI state conflict resolution is core functionality, not optional polish — it must ship alongside the initial feature.
Common Follow-Up Questions
- Walk me through what a sparse GSI actually buys you over a normal GSI.
- How did you pick TTL values for the AppSync cache without serving stale cart or pricing data?
- What would you do differently on the optimistic-UI rollout, in hindsight?