Back to Portfolio

Infrastructure & AI Engineer

Current

🏢 ACCENTURE · Bangalore, India · Full-time

Apr 2022 – Present

Leading the architecture, deployment, and optimization of enterprise-grade AI platforms, LLMOps pipelines, and cloud infrastructure operations. Bridging the gap between machine learning applications and scalable production systems.

🛠️ Live Interactive Enterprise CI/CD Pipeline Flow Active Simulator
📦
Code Push
Triggers build
⚙️
Jenkins Build
Checkout & compile
🛡️
Security Scan
Trivy & SonarQube
🐳
Docker Build
Package image
☸️
EKS Deploy
Helm rollout
📦

Step 1: Code Push (Trigger)

A code push to the main branch initiates the multi-branch pipeline workflow via an automated GitHub webhook.

// Core Achievements
  • Agentic AI Architecture: Architected and deployed an Enterprise Agentic AI Operating System integrating LangGraph multi-agent orchestration, RAG pipelines, and tool-calling capabilities across 50+ workflows—reducing manual processing time by 65% and achieving 99.8% uptime SLA.
  • MLOps & LLMOps Platform: Designed and operationalized a cloud-native platform on Azure AKS using ArgoCD GitOps, Helm, and Terraform; automated model promotion pipelines with GitHub Actions, cutting deployment lead time from 3 days to under 2 hours.
  • Enterprise RAG Platform: Built an Enterprise RAG Platform with Weaviate vector database, OpenSearch hybrid search, and LangChain retrieval chains—delivering semantic document intelligence across 5M+ knowledge base records with <200ms P99 query latency.
  • Incident Triage Copilot: Engineered an AI Infrastructure Monitoring Copilot using Claude/GPT-4 function calling + LogicMonitor APIs, reducing MTTR by 58% and enabling natural-language incident triage that auto-generated ServiceNow tickets with root cause hypotheses.
  • Multi-Cloud Orchestration: Implemented multi-cloud Kubernetes platforms (AKS + EKS + GKE) with Helm umbrella charts, Istio service mesh, and OPA Gatekeeper policy enforcement—supporting 200+ microservices across 15 development teams.
  • Full-Stack Observability: Established full-stack observability platform (Prometheus, Grafana, Loki, Tempo, OpenTelemetry) with AI-powered anomaly detection, cutting P1 incidents by 42% and enabling distributed tracing across 30+ services.
  • Voice & Assistant Platform: Delivered WhatsApp + Telegram AI Assistant platform using FastAPI, LangChain agents, and Redis caching—serving 10,000+ daily interactions with sub-500ms response times.
  • VMware Virtualization: Administered VMware vSphere 7.x/8.x environment (ESXi, vCenter, vSAN) across Dell MX7000 and Cisco UCS blade infrastructure—managing 200+ VMs, achieving 40% compute consolidation, and saving $300K annually.
  • SRE & Governance: Led Patch Management and Change Management for 500+ Windows Server instances using WSUS/SCCM and ServiceNow CAB processes, maintaining 99.2% patch compliance.
  • Storage Infrastructure: Designed NetApp SAN/NAS storage architecture with SnapVault/SnapMirror replication—providing 50TB+ enterprise storage with 99.99% availability and RPO/RTO of 15 min/1 hour.
LangGraph Weaviate Kubernetes (AKS/EKS) Terraform ArgoCD Prometheus & Grafana VMware ESXi Python (FastAPI) LogicMonitor ServiceNow