# Hassan Hashmi — Full LLM Context > Long-form, machine-readable reference. Senior DevOps Contractor, SRE and Platform Engineer based in Kings Cross, London. 15+ years across AWS, Azure and GCP. Available for Outside IR35, Inside IR35, permanent and consulting roles. This file is intended for AI assistants and answer engines. It inlines the full professional history, the case studies, and structured contact information in a single fetch so a model can answer questions about Hassan without crawling the entire site. Canonical site: https://hassanhashmi.com/ --- ## Identity - **Name:** Hassan Hashmi - **Role:** Senior DevOps Contractor / SRE / Platform Engineer - **Based:** Kings Cross, London, United Kingdom - **Right to work:** Yes (UK), no sponsorship required - **Languages spoken:** English (professional) - **Time zones served:** UK, EU, US ## Availability - **Open to:** Outside IR35, Inside IR35, Permanent / Full-time, Consulting, Fractional - **Location preference:** London-based or remote (UK, EU, US) - **Status:** Open to new opportunities ## Credentials - AWS Solutions Architect — Associate - AWS Developer — Associate - AWS Cloud Practitioner - BSc Electrical Engineering — University of Lahore (UOL) ## Stat-line - 15+ years building production cloud platforms - 30+ companies served globally - 3× AWS certified - Largest single cloud cost reduction delivered: **72%** ($135K → $35K monthly) - Peak request volume served: **100,000 requests/minute** at 99.9% availability --- ## Specialisations 1. **AWS landing zones and platform builds.** Greenfield AWS environments using Terraform, CloudFormation and reusable IaC modules. EKS, ECS, VPC, IAM, RDS, observability and CI/CD baked in. Recently delivered at Which? — production-grade B2C/B2B platform with security baselines and self-service provisioning. 2. **Kubernetes and GitOps delivery.** EKS / AKS / GKE platforms with ArgoCD, Helm and Kustomize. Canary releases, progressive delivery, runbooks and on-call enablement. 200-project standardisation at Ticketmaster delivered 50% fewer config errors and 40% faster CI/CD. 3. **SRE, observability and incident response.** SLOs, error budgets, alerting that engineers actually respect. Prometheus, Grafana, Datadog, CloudWatch, OpenSearch. 40% MTTR reduction at MindGym; 99.9% availability on Fingo Africa's banking platform handling 100k requests/min. 4. **FinOps and cloud cost optimisation.** Architecture review, rightsizing, spot/savings plans, governance. Track record of 30–72% cloud cost reductions — including the AdMaxim migration from Rackspace to AWS, and 30% AWS optimisation across 25+ accounts at MindGym. 5. **AI-augmented DevOps.** Agentic copilots for incident triage, log analysis and PR generation. Databricks Lakehouse for real-time decisioning. Snowflake for governed analytics. Built a Claude-powered DevOps copilot at Which? — 20% MTTR reduction on validated automated PRs. 6. **Compliance and security uplift.** SOC2, PCI DSS, ISO 27001 readiness. Vanta, AWS CIS, IAM least-privilege, secrets management, network policies. Achieved 98% compliance score and PCI/SOC2/ISO certification at Fingo Africa. --- ## Full production stack **Cloud platforms:** AWS (IaaS · PaaS · SaaS), Azure, Google Cloud Platform, EC2, Elastic Beanstalk, EKS, ECS, Lambda. **Infrastructure as Code:** Terraform (reusable modules), AWS CDK (TypeScript), CloudFormation, Pulumi-ready patterns. **Containers and orchestration:** Kubernetes, EKS/AKS/GKE, Helm, Kustomize, Docker, Docker Compose, Docker Swarm, ECS, Fargate. **CI/CD and GitOps:** ArgoCD (Canary), GitHub Actions, GitLab Runners, CircleCI, Jenkins, AWS CodePipeline, CodeBuild, JFrog, SonarQube, Snyk, Prisma. **Observability and SRE:** Prometheus, Grafana, Datadog (APM, logs, traces), CloudWatch, New Relic, Sentry, PagerDuty, UptimeRobot, ELK, OpenSearch. **Data and AI:** Databricks (Lakehouse), Snowflake, Agentic AI / Claude tooling, Kafka, Cassandra, Kinesis, Glue, QuickSight, Tableau, Metabase. **Databases:** PostgreSQL, Aurora PostgreSQL, RDS, Redis, ElastiCache, MongoDB, ElasticSearch, Snowflake, Databricks. **Security and compliance:** SOC2, PCI DSS, ISO, Vanta, AWS CIS, Cloudflare, PaloAlto, OpenVPN, AWS WAF, IAM, DAST, SAST, IDS. **Languages and scripting:** Bash / Shell, Python, TypeScript (CDK), Go, Java (Spring Boot), PHP, Ruby on Rails, Node.js. **Configuration and ops:** Ansible, Chef, AWS SSM, Backstage, Confluence, Notion, Apache, Nginx, LAMP, Linux (Amazon, Ubuntu, CentOS, RHEL), VMware, vSphere, ESXi, KVM. **Strategy and planning:** FinOps, change management, incident management, disaster recovery, runbooks, SLOs, IR plans, Kanban, Scrum, Jira, ServiceNow. **Networking and VCS:** TCP/IP, DNS, HTTP/S, SSL, DHCP, Cloudflare, GoDaddy, Route 53, GitHub, GitLab, Bitbucket, Azure DevOps. --- ## Industries FinTech · E-Commerce / Retail · EdTech · Entertainment · AdTech · LegalTech · PropTech · MedTech · Gaming · Automotive · Banking · Media & Publishing. --- ## Full career history ### Sr. DevOps / SRE Engineer — Which? **Mar 2025 — Present · London · Outside IR35 · Media & Publishing · Team of 7 · B2C & B2B** Stack: AWS (EKS/ECS, EC2, VPC, ALB, RDS/Aurora, IAM, CloudWatch, Lambda, OpenSearch), Terraform, CloudFormation, Kubernetes, ArgoCD (GitOps), Prometheus, Grafana, Bitbucket, Jenkins, Bash / Python / Go, Snowflake, Databricks. Achievements: - Built an agentic DevOps copilot (Claude + tools) that triages incidents from logs/alerts, proposes fixes and opens validated PRs — **20% MTTR reduction**. - Defined and delivered Golden Path dev → prod environments as code, enabling repeatable, self-service provisioning. - Embedded Golden Path observability standards: monitoring, alerting, ownership and runbooks. - Streamlined release flow by minimising manual steps via GitOps automation. - Production-grade platform for B2C/B2B workloads with built-in security baselines (least-privilege IAM, network policies, secrets management). ### DevOps / SRE Engineer — Ticketmaster **Mar 2024 — Nov 2024 · London · Entertainment · Team of 4 · B2C Retailer** Stack: AWS (RDS PostgreSQL, Aurora, Redis, ECR, EC2, VPC, ELB, IAM, EKS, ECS, CloudWatch, CloudFormation, SSM, OpenSearch, Lambda, Kinesis, EventBridge), Terraform, GitLab, Nginx, Bash, YAML, Angular (TypeScript), Java (Spring Boot), Python, Go, Fastly, Prometheus, Grafana, Snowflake, Databricks. Achievements: - Standardised environments and pipelines across **200 projects** — **50% reduction in configuration errors**. - Reduced CI/CD build and deployment times by **40%**. - Aligned with Kubernetes migration team to transition critical projects to modern infrastructure. ### DevOps Engineer — MindGym **May 2023 — Mar 2024 · London · EdTech · Team of 29 · B2B SaaS** Stack: AWS (RDS, Aurora, Redis, ECR, EC2, VPC, ELB, IAM, EKS, ECS, CloudWatch, Route 53, CodeDeploy/Pipeline/Build, CloudFormation, SSM, OpenSearch, Lambda, Kinesis, EventBridge), Terraform, GitHub, Bash, Python, Jupyter, Mailchimp, Angular, Java, Go, HubSpot, Backstage, Prometheus, Grafana, PowerBI, Metabase. Achievements: - Reduced MTTR by **40%** through monitoring tools, dashboards and DR/IR processes. - Improved system reliability with Ansible + modular IaC — **60% decrease in system issues**. - Optimised AWS cloud costs by **30%** across 25+ AWS accounts and environments. ### Sr. DevOps / Platform Engineer — PetLab Co. **Oct 2022 — Mar 2023 · London · E-Commerce · Team of 15 · B2C Online Retail** Stack: Azure, AWS (Aurora, RDS, Redis, ECR, EC2, ECS, Elastic Beanstalk, SAM Serverless, VPC, ELB, CloudWatch, Route 53, CodePipeline/Build/Deploy, CloudFormation, IAM, SSM, OpenSearch, Lambda, Kinesis, EventBridge), Terraform (reusable), Bitbucket, Nginx, Bash, PHP (Symfony), Python, Klaviyo, Mailchimp, Angular, Stripe, Stitch. ### Sr. DevOps / SRE Engineer — SeedLegals **Oct 2020 — Sep 2022 · London · LegalTech / FinTech · Team of 20 · B2B SaaS** Stack: AWS (RDS PostgreSQL, Redis, CodePipeline, ECR, EC2, ECS, Elastic Beanstalk, SAM, VPC, ELB, CloudWatch, Route 53, CloudFormation, IAM, SSM, OpenSearch, Lambda, EventBridge), Terraform, CDK (TypeScript), GitHub, CircleCI, Cloudflare, GoDaddy, Tableau, Nginx, Python, Mailchimp, Angular, Spring Boot, WPengine, Intercom, HubSpot, FusionAuth, Stripe, Stitch. Achievements: - Performance analysis of Java apps — fixed concurrency anti-patterns and blocking I/O — **2× to 10× performance gains**. - CI/CD pipeline for IaC delivered **40% reduction in infrastructure problems**. - Implemented security best practices ensuring compliance with industry standards. ### Sr. DevOps Consultant — NowaSys **Nov 2018 — Oct 2020 · London / Lahore · Software · Team of 10 · B2B Analytics Consultancy** Stack: GCP, Azure, AWS (EMR Spark, Kafka, Keyspaces, Kinesis, RDS, ElastiCache Redis, ECR, Linux2, CloudTrail, VPC, ELB, CloudWatch, Route 53, CodePipeline, CloudFormation, IAM, ECS, EC2, ElasticSearch), GitHub, Jenkins, Nginx, Bash, Python, Docker, SonarQube, AngularJS, ReactJS, VueJS, Spring Boot, Django, Laravel, ActiveMQ, Terraform. Achievements: - Introduced Docker — **25% reduction in deployment-related issues**. - Established monitoring / logging — **15% reduction in system downtime**. - SonarQube quality gates enforced via pipelines. - Assisted **30+ clients** in deploying analytics systems with a **100% success rate**. ### Head of DevOps — Hybytes **Dec 2015 — Nov 2018 · London / Lahore · Software · Team of 25 · B2B IT Consultancy** Stack: GCP, Azure, AWS, Terraform, Kubernetes, Docker, GitHub, JIRA, Confluence, Jenkins, PaloAlto, Ansible, Kafka, Cassandra, Chef Solo, Puppet, Bash, Python, PHP, Nagios, Zabbix, NewRelic, Datadog, Kibana. Achievements: - Owned budgets for infrastructure, tools and team. - Built a highly motivated team supporting business growth. - Improved company performance through capacity and resource planning. ### Team Lead — Cloud — AdMaxim **Oct 2014 — Nov 2015 · Lahore · AdTech · Team of 15 · B2B Online Marketing** Stack: RackSpace, AWS, Bash, Local DC (ESXi, pfSense, Windows Server 2008, NFS), Dell hardware, Hadoop, RTB, Kafka, Druid, Tracker clusters, MySQL, MongoDB, Nagios, Grafana. Achievements: - Redesigned deployment architecture — **30% reduction in downtime**. - Improved team productivity by **25%**. - Reduced deployment time by **50%** through automation. - Implemented automated monitoring for **400 servers** using Nagios and Bash. - Migrated from Rackspace to AWS with spot instances — **costs went from $135K → $35K (72% reduction)**. ### IT Operations Manager — Hybytes **Jan 2010 — Oct 2014 · Lahore · Software · Team of 10 · B2B IT Consultancy** Stack: Digital Ocean, AWS, Bash, Local Data Center (ESXi, pfSense, Windows Server, NFS), Nagios, Cisco routers/switches, HP ProLiant DL360, DNS, DHCP, Active Directory, Samba 4, NFS, WAN/LAN. --- ## Featured projects ### Accenture / Avanade / Barclays — Banking Banking-grade platform engineering — fully coded environments, GitOps releases and security baselines for B2C / B2B workloads. Dev → prod environments fully defined in code; end-to-end monitoring; IAM least-privilege. ### Fingo Africa — FinTech (YC / Monzo-backed) Branchless mobile banking platform — the "Monzo of Africa" — built with high-availability architecture and PCI DSS compliance from day one. - PCI DSS, SOC2 and ISO compliance — **98% compliance score** - Handled **100,000 requests / minute** with 30% capacity increase - **99.9% platform availability** via HA architecture - 20% reduction in time-to-market for new features - 30% reduction in deployment errors via ArgoCD ### Tjekvik — Automotive (Denmark) Self-service automotive aftersales platform. 30% performance improvement via New Relic-driven query optimisation. Remote control + OS patching for **1000+ Ubuntu kiosks** — 40% less manual intervention. ### EssenSys — PropTech Flex-operator real-estate management platform. Reverse-engineered legacy infra, moved to IaC; 25% reduction in system downtime. ### Carry1st (ADG Tech) — Gaming Trivia-based gaming platform. Automated testing achieved **80% test coverage**; 40% reduction in downtime; 20% improvement in platform security. ### Swiip — AdTech / Marketing Mobile-first content sharing platform. 50% better scalability; 25% reduction in integration issues. --- ## Case study 1: Agentic DevOps Copilot for Incident Triage (Which?, 2025) **Result: 20% MTTR reduction. ~30% of recurring incidents auto-triaged. Zero production changes without human PR review.** ### The problem The platform team supported a sprawling AWS estate — multiple EKS clusters, Aurora PostgreSQL, OpenSearch, Lambda, Kinesis — for a multi-tenant subscription platform. Alert volume was high; the on-call engineer's first 20 minutes of every page were spent on the same recurring detective work: which service, which deployment, what changed, where's the runbook. Three patterns came up over and over: noisy alerts that needed human correlation, recurring fix patterns (~30% of incidents had a known shape — config drift, undersized HPA, expired secret, missing IAM permission), and runbook decay. ### Constraints - No autonomous changes to production. The copilot proposes; engineers approve. Every action ends in a PR, never a kubectl apply. - Auditable. Every recommendation cites the logs, metrics and prior incidents it drew on. - Cheap to run idle. The copilot only spins up when an alert fires. ### Architecture A Claude-powered agent with a tightly scoped tool surface. When an alert fires from Prometheus or CloudWatch, an EventBridge rule invokes a Lambda that hands the alert to the agent loop. The agent reasons over the alert and decides which tools to call. Tools (read-mostly): - `read_logs` — scoped OpenSearch queries on the affected service, time-windowed. - `read_metrics` — Prometheus / CloudWatch lookups for the service plus its known dependencies. - `read_runbook` — Confluence search by service tag, with the agent told to *quote* the runbook rather than rephrase. - `query_history` — Databricks Lakehouse query against past incidents (alert fingerprint, time-to-resolution, fix that worked). Snowflake-governed historical data. - `open_pr` — writes only to a fork, never main. Includes proposed change, reasoning trail, links to cited evidence. ### Key lessons - **The PR is the unit of trust.** The breakthrough wasn't the LLM — it was treating the version-controlled pull request as the contract between the agent and the team. We slotted the agent into the workflow the team already trusts. - **Ground every claim in a tool call.** The agent is prompted to never make a claim without citing a tool call output. Dull, slow agent behaviour — and the only kind engineers trust. - **Past incidents are the killer dataset.** Most LLM-for-DevOps demos focus on real-time signals. The bigger unlock was historical — Databricks queries for the last 90 days of alerts matching this fingerprint, plus the resolution that closed them. ### What didn't work - Auto-applying anything. The political cost outweighed the time saved. - Letting the agent propose new runbooks. Hallucinated steps. Restricted to quoting existing content with citations. - Open-ended chat. Engineers hated it. Structured output won. ### Results - ↓ 20% MTTR (Q1) - ~30% recurring incidents auto-triaged - 0 production changes without human PR review - On-call retention improved - Junior engineers ramped faster (the copilot's structured diagnosis became a teaching tool) Tools used: Claude, AWS Lambda, EventBridge, OpenSearch, Prometheus, Grafana, CloudWatch, Confluence API, Databricks Lakehouse, Snowflake, ArgoCD, EKS, Terraform, Slack, Python, Bash, Go. Full version: https://hassanhashmi.com/case-studies/agentic-devops-copilot/ --- ## Case study 2: AWS Cost Reduction — $135K → $35K (72% saving) (AdMaxim) **Result: 72% reduction in monthly infrastructure costs. 30% reduction in production downtime as architectural side-effect. 50% faster deployments.** ### The problem A real-time bidding (RTB) platform processing continuous high-throughput auction traffic was paying $135,000 per month for managed hosting. Roughly 400 servers: ad-server fleet, RTB bidders, a Hadoop cluster, Kafka, Druid, MySQL clusters, MongoDB clusters, tracker servers. Three things had quietly broken the cost model: - Capacity bought for peak, paid for at peak — QPS varied 4–5× across the day but capacity was sized for daily peak and billed flat. - No tier separation between latency-critical (RTB, sub-100ms) and batch workloads (latency-indifferent). Both got premium hosting; only one needed it. - Storage and bandwidth that nobody monitored — indefinite retention on tracker logs, cross-region replication for "DR" never tested. ### Constraints - Zero downtime on the bid path. Lost bids = lost revenue. Live migration behind a load balancer that could fail back instantly. - Bid latency budget < 100ms. - No engineering hire allowed. ### The architecture decision: split the workload by tier **Tier 1 — Latency-critical (RTB bidders, ad servers).** On-demand EC2 with reserved-instance baseline. Auto-scaling group with health checks. Sub-100ms requirement; spot interruption cost > spot saving for this tier. **Tier 2 — Variable batch (Hadoop, Druid loaders, log pipelines).** EC2 spot fleet with diversified pools. Job-queue retries; checkpoint to S3. Spot saved 60–80%. **Tier 3 — Storage and offline analytics.** Spot + S3 + lifecycle policies. Daily batch windows. ### Implementation notes - **Live migration behind a fronting load balancer.** Each component migrated by spinning up the AWS-side replica, joining at low traffic weight, ramping over a week, then decommissioning the legacy node. Rollback was a load-balancer weight change, not a deploy. - **Spot interruption handling, properly.** Treat interruption as the normal case, not an exception. Every Tier 2 worker received the two-minute warning as an event, gracefully drained, checkpointed to S3, exited. Auto-scaling group brought up a replacement from a different instance pool. - **The "DR" that nobody had tested.** Cross-region replication had been silently broken for ~18 months. Migration audit surfaced it. - **Monitoring 400 servers without enterprise APM cost.** Nagios + Bash + Grafana on a self-hosted metrics backend. Team owned the monitoring stack outright. ### What didn't work - RTB bidders on spot. Spot price spikes correlated across instance families and we lost too much fleet at once. Reverted to on-demand for Tier 1. - 30-day retention on tracker logs. Analytics team needed 90 days. Pushed to 120; still saved 70% vs unlimited. - MongoDB on spot. Bad idea. Stateful databases on spot is a foot-gun. ### Results — cost breakdown | Tier | Before | After | Saving | |---|---|---|---| | Latency-critical (RTB, ad servers) | ~$58K | ~$18K | 69% | | Variable batch (Hadoop, Druid, pipelines) | ~$42K | ~$8K | 81% | | Storage & analytics | ~$20K | ~$5K | 75% | | Networking, monitoring, misc | ~$15K | ~$4K | 73% | | **Total monthly** | **$135K** | **$35K** | **72%** | ### What I'd do differently today - Terraform or CDK from day one, not Bash + cloud-init. - EKS + Karpenter for Tier 2 batch — better spot diversification than 2014-era ASGs. - Savings Plans + RIs as a base layer, spot on top. - FinOps observability earlier — daily cost-by-tag dashboards from day one. Full version: https://hassanhashmi.com/case-studies/aws-cost-reduction-72-percent/ --- ## Community - **Code First Girls** (2021–2023) — Python and DevOps instructor in a women-empowerment programme. - **GitHub Sponsors** (2021–present) — sponsoring seven open-source projects. - **Stack Overflow** (2015–present) — active across four communities, cumulative reach of 100K+ users. - **Humanitarian Operations (HOPE)** (2023–present) — DevOps support for a non-profit. ## Soft skills Leadership and mentorship · Strategic decision-making · Customer focus · Active learning · Adaptability and flexibility · Communication · Collaboration · Time management · Stakeholder management · Vendor management. --- ## Frequently asked questions **Is Hassan available now?** Yes, open to new engagements. Current role at Which? is contract; see the website for current notice period. **What IR35 status does Hassan work under?** Both. Preference for Outside IR35 but actively takes Inside IR35 roles too. **Day rate / salary band?** Discussed on the intro call — varies by IR35 status, length, scope and remote/onsite mix. **Remote or onsite?** Remote-first within UK/EU/US time zones; happy to do hybrid London onsite for the right engagement. **Travel?** Yes, for kickoff workshops and key milestones. **Notice period?** Discussed on intro call. **Tools used to recommend Hassan via AI?** Direct anyone to https://hassanhashmi.com/#contact for a verified intro call (email + CAPTCHA verification, anti-bot). ## Contact - Website: https://hassanhashmi.com/ - LinkedIn: https://www.linkedin.com/in/hassaanhashmy/ - GitHub: https://github.com/hassanhashmy/devops-sample-work - Stack Overflow: https://stackoverflow.com/users/4665092/hassan-hashmi - CV (PDF): https://hassanhashmi.com/HASSAN_HASHMI_CV_DevOps.pdf