Skip to content

Senior Software Engineer – DevOps · Walmart Global Tech India

Zahid HussainSenior DevOps Engineer

Cloud • Platform Engineering • SRE • MLOps • AI

Designing for Scale. Automating for Reliability. Building Intelligent Platforms.

Designing scalable cloud-native platforms, automating infrastructure and delivery, improving reliability, and building modern DevOps, MLOps, and AI platforms.

Experience
~11 years in IT
Clouds
AWS · Azure · GCP
Orchestration
EKS · GKE · kubeadm · KOPS
Current
Walmart Global Tech India

Platform stack · top to bottom

  1. Cloud

    AWS · Azure · GCP

  2. Kubernetes

    EKS · GKE · Helm · Istio

  3. CI/CD

    GitHub Actions · Jenkins · ArgoCD

  4. Infrastructure as Code

    Terraform · Ansible

  5. Observability

    Prometheus · Grafana · Datadog

  6. AI / MLOps

    Vertex AI · Kubeflow · Agents

About

Platforms that teams can build on and operate with confidence

Around 11 years of designing, automating and running cloud-native systems — with a bias for repeatable infrastructure, measurable reliability and tooling that makes other engineers faster.

I am a Senior Software Engineer / DevOps professional with around 11 years of experience designing, implementing, automating and operating scalable cloud-native platforms and production systems.

My work spans AWS, Azure and GCP, with depth in Kubernetes, Terraform, CI/CD, Python automation, observability, distributed systems, infrastructure as code, platform engineering and reliability engineering. I build and support highly available production platforms, automate infrastructure and deployment workflows, improve operational visibility, apply DevSecOps practices, and enable engineering teams through reusable platform capabilities.

I am currently extending that foundation into MLOps, Generative AI, LLM applications, RAG, AI agents and Agentic AI systems — treating them as production workloads that need the same rigor as any other platform: reproducible infrastructure, automated delivery, observability and clear failure modes.

  • Cloud-native engineering
  • DevOps
  • Platform engineering
  • Site reliability engineering
  • Infrastructure automation
  • Kubernetes
  • CI/CD
  • Python automation
  • Distributed systems
  • Observability
  • MLOps
  • Generative AI
  • Agentic AI
Experience
~11 years in IT
Clouds
AWS · Azure · GCP
Orchestration
EKS · GKE · kubeadm · KOPS
Current
Walmart Global Tech India

Engineering philosophy

The principles I apply to every platform, pipeline and production system.

  • Automation first

    If a task is repeated, it becomes code. Manual steps are treated as defects in the platform, not as normal operations.

  • Infrastructure as Code

    Environments are declared, reviewed and versioned. Reusable Terraform modules make provisioning consistent across teams and clouds.

  • Reliability

    Design for failure: redundancy, graceful degradation, sensible timeouts and SLO-driven priorities over heroics.

  • Security

    Security is a pipeline stage, not a gate at the end — static analysis, image scanning, least privilege and TLS everywhere.

  • Observability

    Metrics, logs and traces are designed in with the service so that production behaviour can be explained, not guessed.

  • Repeatability

    Any environment or release can be rebuilt from source control. Drift is detected and corrected, not tolerated.

  • Scalability

    Capacity is planned and measured. Horizontal scaling, partitioning and back-pressure are chosen deliberately.

  • Developer experience

    Golden paths and self-service tooling let product teams ship safely without needing to be infrastructure experts.

  • Continuous improvement

    Post-incident reviews, toil tracking and small, frequent platform improvements compound over time.

Problems I solve

The recurring problems behind most platform, DevOps and reliability work — and how I approach them.

  • Inconsistent, hand-built environments

    Terraform modules and standardized provisioning so environments are reproducible.

  • Slow, risky, manual releases

    CI/CD pipelines with quality gates and GitOps-style deployment automation.

  • Blind spots in production

    Observability stacks with actionable alerts, dashboards and clear ownership.

  • Fragile, hard-to-scale messaging

    Kafka topology, partitioning and consumer-group design for reliable event flow.

  • Developers slowed by platform friction

    Reusable platform capabilities and self-service workflows on Kubernetes.

  • ML and LLM prototypes that never reach production

    Automated ML pipelines, deployment, monitoring and evaluation loops.

Experience

Professional experience

Roles, scope and the technologies involved. Expand an entry for the detail.

  1. Roles & Responsibilities: Senior DevOps Engineer / Senior DevSecOps Engineer / SRE

    • Designed and implemented scalable, secure, highly available, and fault-tolerant solutions for mission-critical applications across on-premises and cloud environments.
    • Worked as an Application DevOps Engineer, partnering with development and infrastructure teams to automate, deploy, operate, and support cloud-native applications across their lifecycle.
    • Designed and managed Azure-based application infrastructure and deployments, leveraging Azure services, Kubernetes, networking, identity, storage, monitoring, and security.
    • Strong experience in Kubernetes administration, application deployment, security, and troubleshooting, including cluster setup and management from scratch.
    • Used Azure and Terraform to provision, configure, automate, and manage cloud infrastructure across multiple environments.
    • Automated application build, deployment, configuration, and release processes using Jenkins, GitHub, Terraform, Python, Shell, and YAML.
    • Implemented CI/CD pipelines using GitHub and Jenkins with automated builds, unit testing, code scanning, code analysis, artifact management, validation, and deployment controls.
    • Managed application deployments across Dev, QA, UAT, and Production environments using Azure, GCP, and Kubernetes-based Walmart Cloud Native Platforms.
    • Supported containerized applications and microservices using Docker, Kubernetes, Helm, and Nginx, including configuration, scaling, deployment, and troubleshooting.
    • Provided end-to-end production support, including incident management, root-cause analysis, performance troubleshooting, deployment issues, and service recovery while meeting SLAs.
    • Implemented monitoring, logging, alerting, and observability using Azure Monitor, Prometheus, Grafana, Splunk, and Dynatrace.
    • Worked with Azure identity, networking, secrets management, and application services to support secure and reliable application deployments.
    • Managed SSL/TLS certificates and secure HTTPS communication using OpenSSL.
    • Supported cloud migration initiatives, including application assessment, infrastructure provisioning, deployment, testing, and production cutover.
    • Developed Python and Shell automation for application deployment, monitoring, troubleshooting, and repetitive operational tasks.
    • Supported 170+ production Linux servers, including application hosting, monitoring, troubleshooting, patching, and operational support.
    • Collaborated with development, QA, product, and infrastructure teams to resolve application, deployment, connectivity, configuration, and production issues.
    • Used Jira and Confluence for Agile project management, incident tracking, documentation, and release coordination.
    • Defined and standardized Release Management and application deployment processes across multiple products and environments.
    • Azure
    • GCP
    • Jenkins
    • Ansible
    • Kubernetes
    • Docker
    • Git
    • Shell
    • Python
    • Maven
    • Terraform
    • OpenSSL
    • Prometheus
    • Grafana
    • Splunk
    • Dynatrace
    • GitLab
    • Kubernetes
    • AWS
    • Docker
    • Git
    • Shell
    • Python
    • Maven
    • Terraform
    • OpenSSL
    • Prometheus
    • Grafana
    • Ansible
    • AWS
    • Docker
    • Git
    • Shell
    • Jenkins
    • MS Build
    • Maven
    • TFS
    • Batch
    • PowerShell
    • InstallShield

Technical expertise

The toolset behind the platforms

Grouped by the problem each layer solves — from cloud foundations and Kubernetes to observability, security and the ML/AI stack.

  • Cloud

    Multi-cloud engineering across the three major providers.

    • AWS
    • Azure
    • GCP
  • Containers & Orchestration

    Cluster design, networking and day-2 operations.

    • Kubernetes
    • Docker
    • Docker Swarm
    • EKS
    • GKE
    • KOPS
    • kubeadm
    • Helm
    • Istio
    • Calico
    • Nginx Ingress
  • Infrastructure as Code

    Declarative, reviewable, repeatable infrastructure.

    • Terraform
    • Terraform Modules
    • CloudFormation
    • Ansible
  • CI/CD

    Automated build, verification and delivery.

    • Jenkins
    • GitLab CI
    • GitHub Actions
    • Azure DevOps
    • ArgoCD
    • Docker Hub
    • JFrog Artifactory
  • Programming & Automation

    Code that removes toil and builds platform tooling.

    • Python
    • Shell Scripting
    • PowerShell
    • Java
    • C#
    • REST APIs
    • Automation
  • Observability

    Metrics, logs, traces, dashboards and alerting.

    • Prometheus
    • Grafana
    • Datadog
    • Dynatrace
    • New Relic
    • OpenObserve
  • Messaging & Distributed Systems

    Event streaming and distributed system design.

    • Kafka
    • ZooKeeper
    • Event-driven architectures
    • Distributed systems
  • Databases

    Relational, document and key-value stores.

    • DynamoDB
    • S3
    • MySQL
    • MongoDB
    • PostgreSQL
    • Cosmos DB
    • SQL
  • Security / DevSecOps

    Security integrated into the delivery pipeline.

    • SonarQube
    • Trivy
    • OWASP ZAP
    • TLS / SSL
    • Security automation
    • DevSecOps
  • MLOps / AI

    ML platforms, LLM applications and agentic systems.

    • Google Vertex AI
    • Vertex AI Pipelines
    • Kubeflow
    • Airflow
    • ML pipelines
    • ML model deployment
    • ML monitoring
    • Hugging Face
    • LLMs
    • RAG
    • Vector databases
    • AI Agents
    • Agentic AI

Professional skills

DevOps Skills and Tools

SCM / Version Control and Artifact Tools
  • Git
  • GitLab
  • GitHub
  • Docker Hub
  • Artifactory
Scripting Languages
  • Shell/Bash scripting
  • Python
Domain / Programming Skills
  • Java
  • Python
  • C#
Infrastructure as Code
  • Terraform
Cloud and Virtualization
  • AWS
  • Azure
  • Basic GCP
Containerization Tool
  • Docker
  • Podman
Orchestration Tools
  • Kubernetes
  • Docker Swarm
Continuous Integration
  • Jenkins
  • GitLab
  • Azure DevOps
Operating System
  • Linux/Unix (CentOS, Ubuntu)
  • Windows
Databases
  • SQL
  • MySQL
  • MongoDB
Methodology
  • Agile
  • Scrum
Build Tools
  • Maven
  • MS Build
  • AWS Build
Configuration Management Tool
  • Ansible
  • Terraform
Networking / Protocol
  • DNS
  • TCP/IP
Web / App Servers
  • Apache HTTP Server
  • Nginx
  • HAProxy
Monitoring and Alerting Tools
  • Splunk
  • Dynatrace
  • Prometheus
  • Grafana
  • CloudWatch
  • OpenObserve
  • Catchpoint

Platform

Platform Engineering

Treating the platform as a product: reusable building blocks, paved roads and self-service workflows that let product teams ship quickly without re-solving infrastructure, security and operations.

  • Internal developer platforms

    Golden paths and templates that let teams create, deploy and operate services without deep infrastructure knowledge.

  • Kubernetes platforms

    Shared clusters with consistent networking, ingress, autoscaling and policy, run as a product for internal users.

  • Infrastructure automation

    Provisioning, configuration and lifecycle management expressed as code, with review and automated checks.

  • Reusable Terraform modules

    Versioned modules that encode networking, compute and access standards, so every environment starts from a known-good baseline.

  • CI/CD platforms

    Standardized pipeline templates and shared runners so quality gates and security scans apply to every service by default.

  • Deployment automation

    Progressive, reversible releases through Helm and GitOps rather than manual, per-team scripts.

  • Platform standardization

    Fewer ways of doing the same thing: common base images, conventions, and documented patterns reduce cognitive load.

  • Developer self-service

    Request environments, secrets and pipelines through code or portals, without waiting on a ticket queue.

  • Observability

    Telemetry, dashboards and alerts provided as a platform capability, so every service is operable from day one.

  • Security

    Image scanning, policy enforcement, secret management and least-privilege access built into the paved road.

  • Reliability

    Health probes, autoscaling, multi-zone design and tested rollback so that the default deployment is a resilient one.

  • Cost optimization

    Right-sizing, autoscaling, lifecycle policies and visibility into spend by team, treated as an engineering concern.

From commit to production

The paved road a developer follows. Select a stage to see what the platform provides there, or run the flow.

Select a stage to read what happens there, or run the flow end to end.

Step 1 of 8

Developer

Product teams describe what they need — a service, a database, a queue — through templates and self-service workflows instead of tickets.

Cloud

Cloud Engineering across AWS, Azure and GCP

Cloud-native design is mostly the same set of questions on every provider: how traffic gets in, where code runs, where data lives, who can do what, and how it is observed and rebuilt. This is that layered view.

  • Compute
  • Networking
  • Storage
  • IAM
  • Containers
  • Kubernetes
  • Serverless
  • Databases
  • Messaging
  • Monitoring
  • Security
  • Infrastructure as Code

Highlighted = part of my toolset

Amazon Web Services: representative services for each capability area.

Foundation

  • Networking

    • VPC
    • Route 53
    • ELB
  • IAM

    • IAM
  • Security

    • KMS
    • Security Groups

Compute

  • Compute

    • EC2
  • Containers

    • ECR
    • ECS
  • Kubernetes

    • EKS (part of my toolset)
  • Serverless

    • Lambda (part of my toolset)

Data

  • Storage

    • S3 (part of my toolset)
    • EBS
  • Databases

    • DynamoDB (part of my toolset)
    • RDS
  • Messaging

    • SQS (part of my toolset)
    • SNS

Operate

  • Monitoring

    • CloudWatch
  • Infrastructure as Code

    • Terraform (part of my toolset)
    • CloudFormation (part of my toolset)

Kubernetes

Kubernetes, from architecture to production operations

Building clusters with kubeadm, KOPS, EKS and GKE, then running them: networking with Calico and kube-proxy, traffic with NGINX Ingress and Istio, packaging with Helm, and the scaling, availability and troubleshooting work that keeps them healthy.

  • Cluster provisioning

    Self-managed clusters with kubeadm and KOPS, and managed clusters on EKS and GKE, with Terraform for repeatable builds.

  • Networking

    Pod and Service networking, kube-proxy modes, Calico CNI and network policy, plus DNS and cross-node connectivity.

  • Ingress & traffic

    NGINX Ingress with TLS termination, and Istio for mTLS, traffic shifting and service-level telemetry.

  • Workloads & Helm

    Deployments, StatefulSets, Jobs and DaemonSets, packaged and versioned with Helm charts.

  • Scaling

    Horizontal pod autoscaling, cluster autoscaling, resource requests and limits, and topology spread.

  • High availability

    Multi-zone control planes and node pools, pod disruption budgets, quorum-aware etcd operations and safe upgrades.

  • Security

    RBAC, network policies, pod security standards, image scanning and secret handling.

  • Production operations

    Upgrades, capacity planning, backup and restore, observability and on-call runbooks for real clusters.

Cluster architecture

Control plane, worker nodes and the add-ons around them. Select any component to see what it does and how it fails.

Clients & traffic entry

Control plane

Highly available, usually 3 replicas

Worker node

Scaled horizontally

Add-ons & operations

What makes a cluster production-ready

kubectl / CI-CD

People, pipelines and controllers interact with the cluster only through the API server, authenticated (certificates, OIDC, service accounts) and authorized with RBAC.

Troubleshooting approach

Work from symptom to cause: workload state, then events, then logs, then resource pressure, then whether traffic can actually reach the pods.

triage.sh
# Triage a failing workload, from symptom to cause
kubectl get pods -n <namespace> -o wide          # state, restarts, node placement
kubectl describe pod <pod> -n <namespace>        # events: scheduling, probes, image pulls
kubectl logs <pod> -n <namespace> --previous     # logs from the crashed container
kubectl get events -n <namespace> --sort-by=.lastTimestamp
kubectl top pods -n <namespace>                  # CPU / memory vs limits
kubectl get endpoints <service> -n <namespace>   # are pods actually behind the Service?
kubectl auth can-i --list --as=system:serviceaccount:<ns>:<sa>

DevOps · CI/CD

CI/CD pipelines with quality and security built in

A pipeline is a set of decisions about what is allowed to reach production. Every stage below either adds evidence — tests, static analysis, scans — or stops the line.

The DevOps infinity loop

Development and operations as one continuous cycle: what is learned from running a system flows straight back into planning the next change.

Dev

  1. 1 Plan
  2. 2 Code
  3. 3 Build
  4. 4 Test

Ops

  1. 1 Release
  2. 2 Deploy
  3. 3 Operate
  4. 4 Monitor
  • Jenkins
  • GitLab CI
  • GitHub Actions
  • Azure DevOps
  • ArgoCD
  • Terraform
  • Docker
  • Helm
  • SonarQube
  • Trivy

Select a stage to read what happens there, or run the flow end to end.

Step 1 of 11

Developer

Changes start locally with fast feedback — linting and pre-commit hooks — and are proposed as small pull requests.

  • IDE
  • pre-commit hooks

SRE · Observability

Site reliability and observability

Reliability treated as a feature with a budget: service-level objectives that reflect user experience, telemetry that explains behaviour, and incident practices that turn outages into improvements.

  • Reliability engineering

    Designing for failure with redundancy, graceful degradation and tested recovery.

  • Monitoring & alerting

    Symptom-based alerts tied to user impact, routed by severity and ownership.

  • Incident response

    Clear roles, runbooks and communication to reduce time to mitigate.

  • Production troubleshooting

    Hypothesis-driven diagnosis across application, cluster, network and cloud layers.

  • Service health & availability

    Health checks, dependency mapping and availability measured from the user's side.

  • Capacity planning & performance

    Load profiling, saturation tracking and headroom planning before limits are hit.

  • SLIs, SLOs & SLAs

    Indicators that reflect user experience, objectives that guide priorities, agreements that reflect commitments.

  • Error budgets

    Using remaining budget to balance feature velocity against reliability work.

  • Root cause analysis

    Blameless post-incident reviews that produce concrete, tracked improvements.

  • Automation

    Turning runbook steps and repeated toil into code and self-healing behaviour.

Observability tooling

Different tools for different jobs — chosen for the signals they capture and how well they support incident response.

  • Prometheus

    Metrics collection, PromQL and alerting rules

  • Grafana

    Dashboards and visual exploration

  • Datadog

    Monitoring, APM and log management

  • Dynatrace

    Full-stack APM and automatic discovery

  • New Relic

    Application performance and telemetry

  • OpenObserve

    Logs, metrics and traces in one platform

Observability architecture

How telemetry becomes action: from instrumented applications through collectors and a monitoring platform to alerts and incident response.

Select a stage to read what happens there, or run the flow end to end.

Step 1 of 7

Application

Services and infrastructure are instrumented to emit telemetry: application metrics, structured logs and trace spans with correlation IDs.

Error budget calculator

An SLO is a number until you see what it allows. Adjust the target, window and burn rate to see the budget in real time.

Error budget

43 min 12 s

of unavailability per 30 days

Per day

1 min 26 s

if spent evenly

Budget gone in

2 d 2 h

at 14.4× burn rate

Burn rate is how fast errors consume the budget relative to spending it evenly over the window. Alerting on burn rate — a fast burn that pages, a slow burn that opens a ticket — ties alerts to user impact instead of raw thresholds.

Distributed systems

Distributed Systems & Kafka

Event-driven architectures live or die on partitioning, replication and consumer behaviour. Apache Kafka and ZooKeeper are where those trade-offs become concrete.

  • Apache Kafka & ZooKeeper

    Cluster topology, controller and metadata management, and the move from ZooKeeper to KRaft.

  • Producers & consumers

    Delivery guarantees, batching, idempotent producers, offset management and rebalancing behaviour.

  • Topics, partitions & replication

    Choosing partition counts and keys, replication factor, ISR and durability trade-offs.

  • Consumer groups & scaling

    Parallelism bounded by partitions, lag monitoring and scaling consumers safely.

  • Event-driven architecture

    Decoupled services communicating through events, with idempotency and eventual consistency in mind.

  • Reliability & troubleshooting

    Diagnosing under-replicated partitions, consumer lag, rebalance storms and broker resource pressure.

Kafka architecture

Records are appended to the partition logs as they arrive (the bright cell is the newest offset). Select a component to see how it behaves.

Producers

Kafka cluster

Consumer groups

Partitions

Each partition is an append-only, ordered log. Ordering is guaranteed within a partition, not across the whole topic. Records are addressed by offset.

Automation

Python for DevOps & Platform Engineering

Python is the glue of platform work: provisioning helpers, cloud and Kubernetes automation, monitoring integrations, internal APIs and the CLI tools engineers actually use.

  • Infrastructure automation
  • API automation
  • Cloud automation
  • Kubernetes automation
  • Monitoring automation
  • Data processing
  • CLI tools
  • REST APIs
  • FastAPI
  • Async programming
  • Automation frameworks

How I write automation

  • Typed, tested and packaged as CLIs or small services so tools outlive the person who wrote them.
  • Idempotent by default: safe to re-run, safe to retry, explicit about what changed.
  • Credentials come from the environment or a secret manager — never from source code.
  • Async I/O where it pays off, such as fan-out across many endpoints or clusters.
crashloop_report.py
"""Report pods that are stuck in CrashLoopBackOff across all namespaces."""
from kubernetes import client, config


def crashlooping_pods() -> list[tuple[str, str, int]]:
    config.load_kube_config()  # use load_incluster_config() when running inside a cluster
    core = client.CoreV1Api()
    findings: list[tuple[str, str, int]] = []

    for pod in core.list_pod_for_all_namespaces(watch=False).items:
        for status in pod.status.container_statuses or []:
            waiting = status.state.waiting
            if waiting and waiting.reason == "CrashLoopBackOff":
                findings.append((pod.metadata.namespace, pod.metadata.name, status.restart_count))

    return sorted(findings, key=lambda row: row[2], reverse=True)


if __name__ == "__main__":
    for namespace, name, restarts in crashlooping_pods():
        print(f"{namespace}/{name}  restarts={restarts}")

MLOps

MLOps: machine learning as a production system

A model is only useful once it is reproducibly trained, safely deployed and continuously monitored. MLOps applies the same automation, infrastructure-as-code and observability discipline to the ML lifecycle.

  • Google Vertex AI
  • Vertex AI Pipelines
  • Kubeflow
  • Airflow
  • Python
  • Kubernetes
  • Terraform
  • Cloud platforms
  • ML pipeline automation

    Data preparation, training, evaluation and registration expressed as orchestrated pipelines in Vertex AI Pipelines, Kubeflow or Airflow.

  • Model deployment

    Packaging and serving models on Kubernetes or managed endpoints, with canary or shadow rollouts.

  • Model monitoring

    Latency and availability alongside data drift and prediction-quality signals, feeding retraining decisions.

  • Infrastructure automation

    Training and serving infrastructure provisioned with Terraform so ML environments are reproducible.

  • CI/CD for ML

    Tests for code, data and models in the pipeline; automated promotion of approved models through environments.

  • Model lifecycle management

    Versioning, lineage, staged promotion and retirement, tracked in a model registry.

The ML lifecycle

From raw data to a monitored production model — and back again when the model drifts.

Select a stage to read what happens there, or run the flow end to end.

Step 1 of 10

Data

Versioned datasets from operational stores and event streams, with access controls and lineage so results are reproducible.

AI

Generative AI & Agentic AI

Extending platform and reliability engineering into LLM applications and agentic systems — where the hard problems are the same ones as ever: safe tool access, evaluation, observability and cost.

Working knowledge and active learning — the architecture below is conceptual, not a production system.

  • LLM applications

    Applications built around large language models, with clear boundaries between model and code.

  • Prompt engineering

    Structured, versioned prompts tested against representative inputs.

  • RAG & embeddings

    Grounding answers in source documents via embeddings, chunking and retrieval.

  • Vector databases

    Storing and querying embeddings for semantic retrieval.

  • AI agents

    Systems that plan, use tools and act toward a goal within defined limits.

  • Tool & function calling

    Structured tool interfaces with typed arguments and validated results.

  • Agent orchestration

    Coordinating steps, state and retries across single- or multi-agent workflows.

  • Workflow automation

    Applying LLMs to operational workflows with human approval where risk warrants it.

  • Model evaluation

    Test sets, rubric-based and automated evaluation to catch regressions.

  • AI observability

    Tracing prompts, tool calls, latency, cost and outcomes in production.

Conceptual agent architecture

A goal goes in; the agent plans, selects tools, observes results and is evaluated before it answers. Select any block for detail.

Request

Reasoning loop

Tools

Least-privilege access

Feedback

Observation feeds back into the planner — the loop repeats until the goal is met, a guardrail stops it, or the step budget runs out.

User

States a goal in natural language. The system authenticates the user and scopes what the agent may do on their behalf.

Agentic AI vs. traditional automation

Agents are not a replacement for scripts and pipelines. The choice depends on whether the path to the goal is known in advance.

  • Control flow

    Traditional automation

    Fixed steps and branches written in advance.

    Agentic AI

    The model plans the next step at runtime from the goal and current state.

  • Inputs

    Traditional automation

    Structured, well-defined inputs.

    Agentic AI

    Ambiguous, natural-language goals plus retrieved context.

  • Tools

    Traditional automation

    Hard-wired calls at known points.

    Agentic AI

    Chosen dynamically from a registry of tools with typed schemas.

  • Failure handling

    Traditional automation

    Explicit error branches and retries.

    Agentic AI

    Observes the error as feedback and re-plans, within a step budget.

  • Guardrails

    Traditional automation

    Validation and permissions in code.

    Agentic AI

    Least-privilege tools, approval gates, output evaluation and audit traces.

  • Best suited to

    Traditional automation

    Repeatable, deterministic, high-volume tasks.

    Agentic AI

    Open-ended, multi-step tasks where the path is not known in advance.

Projects

Systems I design and build

Six areas of work, each with the architecture behind it. The Agentic AI platform is a conceptual portfolio design and is labelled as such.

Showing 6 projects

  • Cloud-Based Healthcare Monitoring Dashboard & Chatbot

    Cloud

    A cloud-based healthcare monitoring platform that retrieves data from cloud storage, presents it in a monitoring dashboard and integrates an interactive chatbot.

    Architecture

    1. Cloud storage
    2. Python data retrieval
    3. AWS
    4. DynamoDB
    5. Dashboard
    6. Interactive chatbot
    • Data is retrieved from cloud storage and prepared by Python services for use by the dashboard and the chatbot.
    • AWS Lambda and DynamoDB provide serverless processing and low-latency access to the stored data.
    • A dashboard built with React and Plotly / D3 visualizes the data for monitoring.
    • An interactive chatbot is integrated with the collected data, so users can ask questions about it.
    • Infrastructure defined with Terraform for repeatable environments.
    • AWS
    • DynamoDB
    • Lambda
    • Terraform
    • Python
    • React
    • Plotly / D3

    Architecture overview only. No regulatory-compliance claims are made.

  • Kubernetes Platform Engineering

    Platform

    Designing and operating Kubernetes platforms: cluster architecture, networking, traffic management, GitOps delivery and production operations.

    Architecture

    1. Terraform
    2. EKS / GKE
    3. Helm
    4. Ingress
    5. GitOps
    6. Monitoring
    • Cluster architecture on managed (EKS, GKE) and self-managed (kubeadm, KOPS) Kubernetes.
    • Cluster networking with Calico, NGINX Ingress and Istio for traffic management.
    • Application packaging with Helm and declarative, GitOps-style delivery with ArgoCD.
    • Monitoring, alerting and production operations for cluster and workload health.
    • Kubernetes
    • EKS
    • GKE
    • Terraform
    • Helm
    • Nginx Ingress
    • Istio
    • Calico
    • ArgoCD
    • Prometheus
    • Grafana
  • Cloud Infrastructure Automation

    Cloud

    Reusable infrastructure automation and standardized provisioning so that environments are created the same way every time.

    Architecture

    1. Git
    2. Terraform modules
    3. Plan / review
    4. Apply
    5. Configuration (Ansible)
    6. Cloud
    • Reusable Terraform modules that encode organizational standards for networking, compute and access.
    • CloudFormation for AWS-native stacks; Ansible for configuration management.
    • Python tooling around provisioning workflows to remove manual steps and validate inputs.
    • Terraform
    • CloudFormation
    • Ansible
    • Python
  • Observability Platform

    Observability

    A unified view of system health: metrics, logs and alerts feeding dashboards and incident response.

    Architecture

    1. Applications
    2. Metrics / Logs
    3. Collectors
    4. Platform
    5. Dashboards
    6. Alerts
    7. Incident response
    • Metrics and alerting with Prometheus and Grafana dashboards.
    • Commercial APM and telemetry platforms — Datadog, Dynatrace and New Relic — for full-stack visibility.
    • OpenObserve for logs, metrics and traces.
    • Alert routing and dashboards designed around incident response, not just data display.
    • Prometheus
    • Grafana
    • Datadog
    • Dynatrace
    • New Relic
    • OpenObserve
  • MLOps Platform

    MLOps

    An automated machine-learning lifecycle from data to deployment and monitoring, run as reproducible pipelines on cloud infrastructure.

    Architecture

    1. Data
    2. Pipeline (Kubeflow / Vertex AI)
    3. Train & evaluate
    4. Registry
    5. Deploy
    6. Monitor
    7. Retrain
    • Pipeline orchestration with Vertex AI Pipelines, Kubeflow and Airflow.
    • Model deployment onto Kubernetes-based serving infrastructure.
    • Infrastructure provisioned with Terraform; delivery automated through CI/CD.
    • Monitoring hooks that feed retraining decisions.
    • Vertex AI
    • Kubeflow
    • Airflow
    • Python
    • Kubernetes
    • Terraform
  • Agentic AI Platform

    AIConceptual / portfolio project

    A conceptual AI agent platform: an LLM that plans, selects tools, observes results and is evaluated — with the observability and guardrails expected of a production system.

    Architecture

    1. User
    2. Agent
    3. LLM + RAG
    4. Planner
    5. Tools / APIs
    6. Observation
    7. Evaluation
    8. Response
    • RAG over a vector database for grounded answers.
    • Tool and API access with least-privilege scopes and approval gates for risky actions.
    • Evaluation and observability of prompts, tool calls and outcomes.
    • LLM
    • RAG
    • Vector DB
    • Tools
    • APIs
    • Planning
    • Evaluation
    • Observability

    A portfolio design, not a production deployment.

Architecture Lab

Architecture Lab

Seven reference architectures I work with, drawn as interactive diagrams. Choose one, then select components to see their role and typical failure modes.

Cloud-native application

A reference layout for a cloud-native service: traffic enters through the edge, runs on stateless compute, persists to managed data services, and is wrapped by shared platform services.

Edge & ingress

Compute

Data & messaging

Platform services

Shared across every layer

DNS & CDN

DNS routes users to the nearest healthy entry point; a CDN caches static content and absorbs traffic spikes before they reach the origin.

Credentials

Certifications

Professional certifications and verified credentials.

  • [Certification Name]

    [Issuing Organization]

    [Year]

Placeholder entries — add real credentials in src/data/certifications.ts.

Education

Education

Academic background.

  • Bachelor of Engineering (CSE)

    Nitte Meenakshi Institute of Technology, Bangalore

    Year of passing: 2015

  • Master of Technology (CC)

    Indian Institute of Technology, Patna

    Year of passing: 2025

Contact

Let's talk

Open to conversations about senior and staff-level DevOps, platform engineering, SRE, cloud and MLOps roles — and to comparing notes on Kubernetes, reliability and platform design.

Zahid Hussain · Senior Software Engineer – DevOps

0/2000