Staff SRE · Platform Engineering · Dubai

Usman Shahid

Classic SRE discipline applied to the agentic-AI era. 11+ years scaling resilient, cost-effective cloud infrastructure, now building production AI Agent platforms with MCP, multi-agent orchestration, and human-in-the-loop approvals.

SCROLL
01

Impact at a glance

11+
Years scaling cloud infrastructure
300+
Microservices migrated to EKS
12 → 1
API gateways consolidated onto Kong
99.99%
Uptime during gateway cutover
02

Primer · open-source project

An unopinionated, batteries-included agent-orchestration platform. Instead of one giant agent with everything crammed into its prompt, Primer orchestrates fleets of small, focused agents, each with a clean working context, runnable on your own hardware.

THE BET

A small model given a clean, purpose-built context can rival a much larger one. A lever on any model, large or small.

Runtime architecture
EDGES ORCHESTRATION CORE PER-TURN REST · Console MCP clients Channels Slack · Telegram · Discord Triggers cron · delay · webhook Worker pool sessions · chats graph runs directed cyclic graphs Park & resume event-driven · frees compute while waiting Human approval gate the risky LLM providers Tools · MCP Workspaces local · container · k8s Collections semantic search Postgres + pgvector
Yielding, event-driven agentsDirected cyclic graphsWorkspaces & sessionsSemantic searchSlack · Telegram · Discord channelsBuilt-in MCP serverHuman approval gatesVersioned harnessesDynamic tool & agent discovery
9
Capabilities, integrated from day one
3
Workspace backends, local · container · k8s
Apache 2.0
Open-source & self-hostable
View Primer on GitHub
03

Engineering for cost

40% EKS cost cut

Cut EKS cluster spend by 40% without compromising reliability, through workload-aware right-sizing and aggressive but safe consolidation.

  • Karpenter consolidation, bin-packing nodes for higher utilization.
  • ARM (Graviton) and Spot migration for compute-heavy workloads.
  • Vertical Pod Autoscaler right-sizing to eliminate over-provisioning.
  • Custom scheduling to densely co-locate compatible workloads.
04

From 12 gateways to 1

At Careem I built a self-service API Gateway platform on Kong, then reliably migrated live public traffic off 12 legacy reverse proxies and gateways onto it, without dropping a request.

Consolidation
12 LEGACY GATEWAYS UNIFIED PLATFORM UPSTREAMS GW 01 GW 02 GW 03 GW 04 GW 05 GW 06 GW 07 GW 08 GW 09 GW 10 GW 11 GW 12 Kong Self-Service Gateway auth · routing · rate-limit · plugins 300+ microservices ~10k RPS sustained public traffic
01
Shadow traffic
Mirror live requests to Kong and compare responses, zero user impact.
02
Per-route canary
Cut over one route at a time behind weighted routing.
03
Weighted shift
Ramp traffic gradually with instant rollback on any regression.
04
Decommission
Retire each legacy proxy once its routes are fully drained.
12 → 1
Gateways consolidated
~10k
Requests per second
99.99%
Uptime during cutover
0
Requests dropped
05

Experience

Jan 2023 – Present
Careem

Staff Site Reliability Engineer

Technical Lead and SRE Architect mentoring a team focused on infrastructure provisioning and developer experience.

  • Designed and built a self-service AI Agent Platform with MCP tool-calling, multi-agent orchestration, and HITL approval flows, ~200K req/day, ~800 tools, ~60 production agents.
  • Slashed EKS cluster costs by 40% via Karpenter consolidation, ARM/Spot migration, VPA, and custom scheduling.
  • Consolidated 12 legacy API gateways (~10k RPS) onto a single Kong gateway with 99.99% uptime during cutover.
  • Built a fully automated provisioning platform with GitOps (Terraform + Terragrunt), new services from days to minutes.
Dec 2020 – Jan 2023
Careem

Senior Site Reliability Engineer

Led migration of 300+ microservices from legacy AWS infrastructure to AWS EKS.

  • Engineered automated provisioning of multiple EKS clusters with Terraform + Terragrunt.
  • Built the observability stack with Prometheus, Thanos, and Grafana, high-cardinality storage, SLO dashboards, cost tracking.
  • Improved deployment safety via Linkerd service mesh with automated canary rollouts through Flagger.
Jul 2019 – Dec 2020
Intech IIS

Infrastructure Architect & Team Lead

Designed the infrastructure of a Cloud IIoT Platform while leading a team of 6 polyglot engineers in a KanBan model.

  • Built a bespoke Kubernetes Ingress Operator for bare-metal clusters using RedHat operator-sdk.
  • Integrated ArgoCD with Keycloak for a secure, auditable GitOps workflow on Helm-based deployments.
  • Implemented Google SRE golden signals on bare-metal Kubernetes using Prometheus.
Jan 2017 – Jun 2019
Intech IIS

Senior Software / DevOps Engineer

Containerized a distributed Industrial IoT platform and built core delivery infrastructure.

  • Built a serverless framework in Python running code on containers via Kubernetes / Marathon / Metronome.
  • Built an OPC → MQTT data pipeline in C# using Avro RPC and REST, plus a StatsD metric aggregator in Go.
  • Built an OpenID Connect Identity Provider with Spring OAuth2 across AD, LDAP, JDBC, and Mongo backends.
Jun 2014 – Jan 2017
Diyatech / Alachisoft

Software Developer · NCache & NosDB

Core developer for NCache Enterprise (distributed cache for .NET) and NosDB Enterprise (NoSQL JSON store).

  • Reimplemented Microsoft's Collections to prevent Large Object Heap (LOH) leakage.
  • Built file-storage-based B+ Tree indexing with single-attribute and compound index support.
  • Built the query execution system using Streams and nested Enumerables in C#.
06

Skills & Stack

AI / ML Ops

PrimerClaude CodeMCPvLLMLlama.cppLLM-DLangchainLanggraph

Observability

PrometheusGrafanaThanosDynatraceOpenTelemetryFluentbitClickhouse

Kubernetes

HelmFlaggerIstioKarpenterOperator-SDKLinkerdKongArgoCD

Cloud

Amazon Web ServicesEKSVPC NetworkingBare MetalMulticluster

SRE / Platform

GitOpsFinOpsSLIs/SLOsTerraformTerragruntIncident Response

Languages

GoPythonJavaC# / .NET
07

From the blog

All posts
08

Selected Writing

Open to relocation

Let's build reliable
platforms together.

Senior IC roles: Staff SRE, Principal Platform Engineer, or Cloud Infrastructure Architect. Visa-sponsored relocation from UAE.