Einbasaharan — Lead Site Reliability Engineer
E

Einbasaharan

$ Lead Site Reliability Engineer

I keep production boring — building reliable, secure, and automated infrastructure across cloud platforms. Kubernetes, observability, incident response, and AI-assisted operations that eliminate toil one script at a time.

What I'm Doing

SECURITY
🛡️

DevSecOps

Designing CI/CD pipelines with security built in — vulnerability scanning for containers, secrets, and Kubernetes, WAF tuning, IAM governance, and compliance guardrails baked into the delivery process.

CLOUD
☁️

Cloud Engineer

Building secure, cost-efficient cloud environments on AWS and GCP — from Landing Zone setup and IAM governance to multi-cloud migration and Graviton optimization.

RELIABILITY
⚙️

SRE

Keeping systems reliably available through Kubernetes platform management, observability stacks, incident response, and eliminating toil with automation.

AUTOMATION
🤖

Infrastructure Automation

Building the tooling that glues everything together — automation scripts in Bash, Go, and Python, Ansible playbooks, and internal services that eliminate repetitive ops work.

AI
🧠

AI & LLM Ops

Bringing AI into operations — LLM-powered incident summarization, RAG chatbots over runbooks and internal docs, and running GPU workloads and self-hosted models on Kubernetes.

AIOPS
📈

AIOps & Observability

Smarter monitoring with anomaly detection, alert noise reduction, and log-pattern clustering — cutting through alert fatigue so humans only get paged when it actually matters.

AI in My Workflow

>_ Incident response copilots — LLMs draft incident timelines, summarize war-room threads, and suggest probable root causes from logs and recent deploys.
>_ Runbook RAG assistant — internal chatbot grounded in our runbooks, postmortems, and architecture docs, so on-call engineers get answers instead of searching wikis at 3 AM.
>_ AI-assisted IaC & code review — Copilot-style tooling for Terraform, Ansible, and Go, with security linting and policy checks keeping generated code honest.
>_ LLM infrastructure — serving self-hosted models (Ollama / vLLM) on Kubernetes with GPU scheduling, autoscaling, and cost controls.

Toolbox

// INFRASTRUCTURE
Kubernetes AWS GCP Terraform Ansible Docker Linux
// CODE & DELIVERY
Go Python Bash ArgoCD GitHub Actions
// OBSERVABILITY
Prometheus Grafana Loki OpenTelemetry
// AI & LLM
Claude / OpenAI APIs RAG Ollama vLLM LangChain Hugging Face MCP GPU on K8s