AWS · Terraform · CI/CD · Observability · GenAI

I build and run the cloud platforms engineering teams depend on.

Cloud Engineer with 5+ years operating production AWS infrastructure, CI/CD platforms, and observability pipelines for 50+ engineering teams in regulated enterprise environments. I ship production GenAI too: Bedrock-backed chatbots and document pipelines that business teams use every day. And I run a production-grade personal cloud at home, with real users, real on-call, and local LLMs. Because I like this work enough to do it off the clock too.

120+ AWS accounts changed via Terraform network rollouts with zero incidents
10x Faster document processing, 5 minutes to 30 seconds, with AWS Bedrock and Textract
~$350K/yr Cloud cost savings from cleanup and rightsizing automation I built
60% to 85% GenAI chatbot accuracy improvement in production, used by 100+ sales reps

What Makes Me Different

Most engineers talk about observability and on-call. I run both at home.

  • 10+ Docker containers behind Cloudflare Tunnel and Tailscale, with Authentik SSO (OIDC, SAML, LDAP) for every service
  • Full observability stack: Prometheus, Grafana, and New Relic monitoring, with PagerDuty paging me when something breaks
  • 95% uptime over 12 months with real daily users who notice when it goes down
  • Local LLMs on Ollama for private AI workflows with no external model providers
pradheep@homelab: ~ $ status --personal-cloud
platformoperational
uptime95% over the last 12 months
services10+ Docker containers
accessCloudflare Tunnel + Tailscale + Authentik SSO
monitoringPrometheus, Grafana, New Relic
alertingPagerDuty
ailocal LLMs via Ollama
on-callme, sole owner and responder

Focus Areas

Platform work, backed by numbers.

Every focus area below comes with a shipped, measurable result from production, not a bullet point of tools I once touched.

CI/CD Platforms

Jenkins and GitHub at enterprise scale

Jenkins administration with Configuration as Code, ECS-backed agent clusters, GitHub Enterprise Server, Artifactory, SonarQube, Prisma, and Checkov.

500+ pipelines migrated solo, untangling cross-account IAM, KMS, S3, and firewall blockers

AWS Infrastructure

Terraform across a multi-account org

ECS, Lambda, VPC, IAM, S3, RDS, EventBridge, SSM, KMS, and Service Control Policies, all managed as code across a regulated AWS organization.

120+ accounts, zero-incident network rollouts

Observability

Logging, monitoring, and alerting

Cribl, OpenTelemetry, Logz.io, New Relic, Splunk, Grafana, Prometheus, and PagerDuty. Centralized pipelines that teams actually adopt.

500GB+/day platform, teams onboarded with minimal changes on their end

GenAI on AWS

Production AI applications

AWS Bedrock, OpenSearch vector databases, and Textract behind secure multi-account infrastructure in HIPAA-audited environments.

Chatbot accuracy 60% to 85% for 100+ sales reps

Selected Wins

Problems I have actually solved.

Killed a multi-month JVM memory leak

Diagnosed a Jenkins metaspace leak through heap dump analysis, tuned JVM arguments, and cleared terabytes of accumulated EFS data. The platform has run without recurrence for 2+ years.

Cut document processing from 5 minutes to 30 seconds

Built an AI document pipeline with AWS Bedrock and Textract that now handles 200+ biopharma documents per day.

Cut GitHub licensing costs with automation

Built Active Directory-integrated automation that suspends inactive users and offboards terminated employees, reducing the license seats needed and eliminating annual manual audits.

Hardened an entire AWS org

Enforced IMDSv2 across all production and non-production accounts through Service Control Policies, closing off instance metadata attacks at the organization level.

Credentials

Certified, published, and battle-tested.

AWS Certified AI Practitioner ISC2 Certified in Cybersecurity Top performer, AWS GameDay incident-response exercise Published in Springer, Advances in Smart System Technologies MS Computer Systems Engineering, Northeastern University