OPERATIONAL · --:--:-- PKT
// syed furqan ali shah full-stack · platform · observability nowshera, pk · utc+5 · remote-ready

I keep 20,000+ machines honest.

Six years building the infrastructure other engineers stand on — in Go and Python. Monitoring agents on 20k+ VMs at 99.9% uptime. Test pipelines that run 70% faster. Multi-tenant SaaS with real isolation, not a tenant_id column and a prayer.

$ uptime --since 2019-01
20,000+
VMs under live monitoring
99.9%
uptime across the fleet
70%
test execution time, 3× parallel throughput
50%
mean time to recovery
30%
metrics ingestion cost at fleet scale
90%
agent upgrade failures — atomic swaps + rollback

Things that run while I sleep.

S-01

nAgent

● live · 20k+ hosts

Self-updating monitoring agent for a GPU cloud fleet

  • A static agent that upgrades itself across 20k+ VMs — atomic binary swaps with rollback cut upgrade failures by 90%. No SSH loops, no config drift.
  • Killed thundering-herd metric spikes with jittered scheduling and cardinality filtering — 30% off the ingestion bill.
  • GPU detection, TLS 1.3 enforced, structured logs. MTTR halved because the data is actually usable at 3 a.m.
GoPrometheusTLS 1.3fleet ops
S-02

Hyper Sentry

● live · CI backbone

Test orchestration platform

  • Cut manual test execution 70% and tripled parallel throughput — regression cycles that took an afternoon now finish before coffee.
  • Pass rate up 25%, flakiness down 30%. Engineers trust red builds again.
  • Failure triage time halved; caught 10% cost regressions before they hit the invoice.
Pythondistributed execCI/CD
S-03

Ripple

● live · multi-tenant

White-label multi-tenant SaaS portal

  • Custom domains per tenant, strict data isolation, role-based access — enterprise onboarding without an enterprise timeline.
  • Customer analytics across 8 pages with 7 reusable chart components; KPI loads pre-warmed so dashboards open hot.
full-stackRBACanalyticscustom domains
S-04

InfraInsight + Infra-Agent

● live · internal platform

Unified insights platform & telemetry pipeline for Hyperstack

  • One place for contracts, resources, and reporting — replaced a scatter of spreadsheets and tribal knowledge.
  • Telemetry service aggregating VM and environment data with validation and encryption end to end.
  • Scheduled collections, repeatable deployments, fewer humans doing robot work.
PythonGotelemetryautomation
2024-03 → now

NexGen Cloud / Full-stack Developer

GPU cloud infrastructure (Hyperstack) — the machines behind AI training runs.

  • Own the monitoring, telemetry, and insights layer for a 20k+ VM GPU fleet.
  • Built the test orchestration platform the team's release confidence rests on.
  • Shipped the white-label portal that turned single-tenant software into a product line.
2019-01 → 2024-03

Wanclouds Inc / Software Engineer, Python

Multi-cloud migration and disaster recovery tooling.

  • Led an LLM-powered natural-language interface for REST APIs — before it was fashionable.
  • Built disaster recovery for cross-region VPC and Kubernetes backup/restore — the software you hope to never need, tested like you will.
  • Wrote the translation layer that moves VPCs and Kubernetes constructs between clouds, and a workflow orchestration engine adopted across product teams.
2014-09 → 2018-08

UET Peshawar / B.S. Computer Software Engineering

Mardan campus. Where the recursion started.

build

  • Go
  • Python
  • TypeScript / React
  • REST API design

run

  • Kubernetes
  • Docker
  • multi-cloud (IBM, GPU cloud)
  • CI/CD pipelines

watch

  • Prometheus-style metrics
  • structured logging
  • alerting & MTTR discipline
  • cost observability

opinions

  • rollback is a feature, not an apology
  • flaky tests are bugs with better PR
  • a dashboard nobody reads is a screensaver
  • boring deploys are the goal

Hiring for platform, infra, or observability?
Let's talk about what keeps breaking.

Response SLA: same day. Faster than my alerts page.