Home Services Skills Projects About Experience Contact
Docker
Kubernetes
Terraform
AWS
Ansible
Prometheus
Grafana
Python
Linux
GitHub

πŸ‘‹ Hi, I am Shiv Prakash

I am a

Having 7+ years of experience building and operating production infrastructure at scale. From Kubernetes and Terraform to CI/CD and Cloud, I focus on keeping systems reliable, scalable, and efficient. I enjoy solving complex infrastructure challenges, automating workflows, and building things that make engineers’ lives easier. Previously worked in ML/NLP, still a Python enthusiast, and always curious about new technologies.

Services & Expertise

End-to-end infrastructure solutions β€” from design and deployment to monitoring and optimization.

Infrastructure as Code

Designing and managing infrastructure using Terraform, Ansible, and CloudFormation for reproducible, version-controlled environments.

CI/CD Pipelines

Building and optimizing continuous integration and delivery pipelines with GitHub Actions, Jenkins, GitLab CI, and ArgoCD.

Monitoring & Observability

Implementing comprehensive monitoring with Prometheus, Grafana, ELK Stack, and Datadog to ensure system reliability and performance.

Cloud Architecture

Designing scalable, secure, and cost-effective cloud architectures on AWS, GCP, and Azure with best practices.

Container Orchestration

Deploying and managing containerized applications with Docker and Kubernetes for scalable microservices architectures.

Site Reliability Engineering

Implementing SRE practices including SLOs, error budgets, incident management, and chaos engineering for maximum uptime.

Skills & Technologies

A broad toolkit built from years of hands-on experience with modern infrastructure and DevOps tooling.

Containers & Orchestration

Kubernetes Docker Nomad Helm Istio ArgoCD

Cloud & IaC

AWS GCP Terraform CloudFormation

CI/CD & Automation

GitHub Actions Jenkins GitLab CI SDLC Microservices

Monitoring & Observability

Prometheus Grafana ELK Stack PagerDuty OpenTelemetry

Languages & Scripting

Python Bash YAML HCL SQL

OS & Networking

Linux Nginx HAProxy DNS TCP/IP VPN Load Balancing

How I Work

Principles and practices that guide my approach to building reliable systems.

How I Build

  • Infrastructure as Code β€” always
  • Immutable infrastructure over mutable
  • Automate everything repeatable
  • GitOps for deployment workflows
  • Least privilege access by default
  • 12-Factor App principles

How I Operate

  • SLOs and error budgets drive decisions
  • Blameless post-mortems
  • Observability over monitoring
  • Chaos engineering for resilience
  • On-call doesn't mean on-edge
  • Toil reduction as a continuous goal

How I Collaborate

  • Clear documentation for every system
  • Runbooks for operational procedures
  • Knowledge sharing over silos
  • DevOps is culture, not a title
  • Feedback loops make teams better
  • Pragmatic over perfect

Featured Projects

A selection of infrastructure and automation projects I've worked on.

K8s Cluster Automation

Automated Kubernetes cluster provisioning and management using Terraform and Ansible, with GitOps-driven deployments via ArgoCD.

Kubernetes Terraform ArgoCD Helm

Observability Platform

Full-stack monitoring and observability platform using Prometheus, Grafana, and Loki with custom dashboards and alerting rules.

Prometheus Grafana Loki OpenTelemetry

CI/CD Pipeline Framework

Reusable CI/CD pipeline templates for GitHub Actions and GitLab CI, featuring automated testing, security scanning, and multi-environment deployments using Argo.

GitHub Actions GitLab CI Argo Bash

Cloud Security Hardening

Automated security compliance and hardening scripts for AWS environments, including IAM policies, VPC configurations, and CIS benchmarks.

AWS SCP Terraform Security

A Bit About Shiv

I'm a DevOps and Site Reliability Engineer who thrives on building systems that are scalable, resilient, and a joy to operate. My approach blends automation-first thinking with a deep understanding of distributed systems.

I believe infrastructure should be treated as code, deployments should be boring, and on-call rotations should be humane. Whether it's architecting a cloud-native platform, setting up observability from scratch, or reducing deployment times from hours to minutes β€” I bring reliability and pragmatism to every challenge.

When I'm not wrangling containers or debugging, you will find me exploring new technologies, travelling to new places, and continuously learning about the ever-evolving cloud-native ecosystem.

0
Countries Explored
0
Years in Production
0
Industries worked in
0
K8s Certified

Experience & Certifications

Building and operating production systems across AI platforms, travel tech, fintech, and semiconductors.

Jan 2024 β€” Present

Site Reliability Engineer

DeepL Cologne, Germany
AI-powered language platform.
AWSOn-PremTerraformCrossplanePythonKubernetesCloudFlareRoute53
  • Designed and implemented Terraform-managed AWS infrastructure with auditable access controls and fine-grained IAM policies for critical production services.
  • Led production infrastructure migrations across AWS services including SES, SQS, SNS, CloudFront, Route53, and RDS with a strong focus on reliability and operational continuity.
  • Improved operational excellence by automating maintenance workflows and reducing manual intervention through reusable infrastructure automation.
  • Contributed to incident management practices by creating operational documentation, improving on-call processes, and onboarding engineering teams into incident response workflows.
  • Reduced engineering toil by developing reusable automation workflows and operational tooling.
  • Reduced engineering toil by developing reusable Claude Skills and automation workflows.
  • Collaborated cross-functionally with engineering and platform teams to improve infrastructure scalability, resiliency, and maintainability.
Oct 2021 β€” Dec 2023

DevOps Software Engineer

trivago Dusseldorf, Germany
Meta search engine for hotels.
AWSGCPKubernetesNomadElasticSearchPrometheusGrafana
  • Designed and operated cloud-native infrastructure across Kubernetes and Nomad environments supporting highly available production workloads at scale.
  • Built and improved CI/CD deployment pipelines to increase release reliability, operational efficiency and stability.
  • Architected scalable microservices using containerization and API-driven infrastructure integrations.
  • Enhanced observability and incident response with Prometheus and Grafana-based monitoring and alerting.
  • Automated infrastructure and operational workflows using scripting and Infrastructure-as-Code practices.
  • Optimized cloud infrastructure costs while maintaining high availability and performance for critical services.
  • Worked closely with engineering teams on platform reliability, operational readiness, and infrastructure improvements in fast-paced production environments.
Jun 2019 β€” Jul 2021

DevOps & Software Development Engineer

AsknBid Tech Bangalore, India
Early stage fintech startup.
AWSEFK StackEnvoyIstioJenkinsKubesphereKubernetes
  • Built and operated Kubernetes infrastructure on AWS EKS for a fintech platform, supporting scalable and resilient cloud-native applications.
  • Implemented service mesh architecture to improve traffic management, service communication, and platform reliability.
  • Designed centralized logging and observability platforms using the EFK stack, Prometheus, and Grafana to improve debugging and operational visibility.
  • Configured and maintained CI/CD pipelines using K8s, Kubesphere, and Jenkins to streamline application delivery.
  • Contributed to infrastructure reliability, security, and automation for customer-facing services.
Jul 2018 β€” May 2019

ML Engineer

Infineon Technologies Bangalore, India
Semiconductor industry.
PythonNLPMachine LearningPandasTensorFlow
  • Trained and deployed ML tools on internal platforms to enable real-time inference and retraining workflows.
  • Developed NLP solutions utilizing LSTM and KNN algorithms for text processing and classification challenges.
  • Built an intelligent error deduplication tool that reduced debugging and response time by 25%.

Licenses & Certifications

Verified

AWS Certified Advanced Networking – Specialty

Amazon Web Services
Issued: Feb 2026
AWS Networking VPC DNS Load Balancing
Show credential
Verified

Certified Kubernetes Administrator (CKA)

The Linux Foundation
Issued: Jan 2026
Kubernetes Cluster Admin Networking Security
Show credential
Verified

Certified Kubernetes Application Developer (CKAD)

The Linux Foundation
Issued: Aug 2025
Kubernetes App Development Helm Pods
Show credential
Verified

Professional Cloud Developer

Google Cloud
Issued: Dec 2023
GCP Cloud Development App Engine
Show credential
Verified

Professional Cloud Security Engineer

Google Cloud
Issued: Jan 2023
GCP Cloud Security IAM Networking
Show credential
Verified

Professional Cloud Architect

Google Cloud
Issued: Dec 2022
GCP Cloud Architecture Infrastructure
Show credential

Let's Connect

Interested in working together or have a question? Feel free to reach out through any of these channels.