Work Experience
5+ years running cloud platforms in regulated production.
At Roche / Foundation Medicine I built and operated the AWS infrastructure, CI/CD platforms, and observability pipelines that 50+ engineering teams depend on every day, in HIPAA-audited environments where mistakes are expensive.
CI/CD Platform Engineering
- Led a solo, year-long Jenkins platform migration to a hardened AWS environment with stricter networking and firewall controls: migrated 500+ pipelines, built Terraform-managed ECS agent clusters, and implemented Jenkins Configuration as Code.
- Debugged and resolved environment-specific failures across the migration, including S3 bucket permissions, firewall policies, and multi-account IAM and KMS access, keeping 500+ pipelines moving without blocking development teams.
- Resolved a multi-month Jenkins JVM metaspace memory leak through heap dump analysis and JVM tuning, cleared terabytes of accumulated EFS data, and added lifecycle management. The platform has run without recurrence for 2+ years.
- Upgraded the Jenkins host OS from Amazon Linux 2 to Amazon Linux 2023, applying hardening standards and updating user-data scripts, package dependencies, and logging pipelines across systems supporting 500+ active pipelines.
- Built credential rotation automation for Jenkins-integrated services, eliminating expired-secret outages and the recurring support tickets that came with them.
Observability and Security
- Designed and deployed a Cribl logging platform with Terraform, ECS, Lambda, and NLB over TLS 1.3, handling 500GB+ of logs per day with autoscaling, onboarding application teams to centralized logging with minimal changes on their end, and routing security events to Splunk for real-time response.
- Migrated all Jenkins and GitHub platform monitoring from New Relic to Logz.io using OpenTelemetry, rebuilding every dashboard and alert with full parity while optimizing metric ingestion for cost.
- Enforced IMDSv2 across the entire AWS organization, production and non-production, through Service Control Policies, hardening instance metadata access org-wide.
- Shipped GitHub Enterprise audit logs to Splunk, delivering compliance-level traceability across the whole GitHub org as a direct IT Security requirement.
- Served on the cloud platform on-call rotation, resolving 180+ incidents within SLA, authoring root cause analyses, and driving preventive fixes that reduced repeat incidents.
AWS Infrastructure and Automation
- Planned and executed Terraform-managed network layer changes across 120+ AWS accounts, including regulated production environments, standardizing routing architecture with zero deployment incidents.
- Built Active Directory-integrated GitHub Enterprise automation to suspend inactive users and offboard terminated employees, reducing license seats and cost and eliminating annual manual audits.
- Designed AWS resource tagging and cleanup automation that raised tagging compliance to 90% and cut monthly cloud spend by 10%.
GenAI Applications and Infrastructure
- Built two production GenAI sales assistant chatbots on AWS Bedrock and OpenSearch vector databases, improving knowledge retrieval for 100+ sales reps and raising response accuracy from 60% at proof of concept to 85% in production through iterative prompt refinement and stakeholder feedback.
- Built an AI document processing application with AWS Bedrock and Textract, cutting processing time from 5 minutes to 30 seconds per file and enabling bulk processing of 200+ biopharma documents per day.
- Designed the secure multi-account AWS infrastructure behind these applications for HIPAA-audited production, using Terraform, IAM roles, encrypted S3, and controlled cross-account access patterns.
- Partnered with stakeholders through UAT, secured approvals, and pushed GenAI applications into production for business users.
AWS
Terraform
Jenkins
GitHub Enterprise
ECS
Lambda
Cribl
OpenTelemetry
Logz.io
Splunk
AWS Bedrock
OpenSearch
Python
Bash
- Architected an automated AWS EBS cleanup pipeline using Jenkins, Python, Bash, and SharePoint integration, eliminating orphaned storage waste and saving about $300K per year.
- Identified and rightsized underutilized EC2 instances with application teams, cutting unnecessary compute spend by about $48K per year.
- Researched migrating team-specific IAM policies to a centralized tag-based ABAC/RBAC model to simplify access management at scale.
- Completed 100% of ServiceNow cloud operations tickets within urgency-based SLA targets.
AWS
Jenkins
Python
Bash
Cost Optimization
Supported students and course operations for technical systems coursework while completing graduate studies in Computer Systems Engineering.
Contributed software engineering work in an enterprise telecom environment before beginning graduate study.
Built early software engineering experience through application development, debugging, and implementation work in a startup environment.
Beyond the Day Job
I also run production infrastructure at home.
A self-hosted personal cloud with 10+ Docker containers, full observability, PagerDuty alerting, and 95% uptime over 12 months, where I am the sole owner and sole on-call responder.