π Hi, I am Shiv Prakash
Having 7+ years of experience building and operating production infrastructure at scale. From Kubernetes and Terraform to CI/CD and Cloud, I focus on keeping systems reliable, scalable, and efficient. I enjoy solving complex infrastructure challenges, automating workflows, and building things that make engineersβ lives easier. Previously worked in ML/NLP, still a Python enthusiast, and always curious about new technologies.
// What I Do
End-to-end infrastructure solutions β from design and deployment to monitoring and optimization.
Designing and managing infrastructure using Terraform, Ansible, and CloudFormation for reproducible, version-controlled environments.
Building and optimizing continuous integration and delivery pipelines with GitHub Actions, Jenkins, GitLab CI, and ArgoCD.
Implementing comprehensive monitoring with Prometheus, Grafana, ELK Stack, and Datadog to ensure system reliability and performance.
Designing scalable, secure, and cost-effective cloud architectures on AWS, GCP, and Azure with best practices.
Deploying and managing containerized applications with Docker and Kubernetes for scalable microservices architectures.
Implementing SRE practices including SLOs, error budgets, incident management, and chaos engineering for maximum uptime.
// Tech Stack
A broad toolkit built from years of hands-on experience with modern infrastructure and DevOps tooling.
// Philosophy
Principles and practices that guide my approach to building reliable systems.
// Portfolio
A selection of infrastructure and automation projects I've worked on.
Automated Kubernetes cluster provisioning and management using Terraform and Ansible, with GitOps-driven deployments via ArgoCD.
Full-stack monitoring and observability platform using Prometheus, Grafana, and Loki with custom dashboards and alerting rules.
Reusable CI/CD pipeline templates for GitHub Actions and GitLab CI, featuring automated testing, security scanning, and multi-environment deployments using Argo.
// About Me
I'm a DevOps and Site Reliability Engineer who thrives on building systems that are scalable, resilient, and a joy to operate. My approach blends automation-first thinking with a deep understanding of distributed systems.
I believe infrastructure should be treated as code, deployments should be boring, and on-call rotations should be humane. Whether it's architecting a cloud-native platform, setting up observability from scratch, or reducing deployment times from hours to minutes β I bring reliability and pragmatism to every challenge.
When I'm not wrangling containers or debugging, you will find me exploring new technologies, travelling to new places, and continuously learning about the ever-evolving cloud-native ecosystem.
// Career
Building and operating production systems across AI platforms, travel tech, fintech, and semiconductors.
// Credentials
// Get in Touch
Interested in working together or have a question? Feel free to reach out through any of these channels.