Step-by-step troubleshooting guides with ready-to-use commands.
Fix AgentCore Identity AccessDenied errors when an agent acts on behalf of a user by diagnosing identity propagation, the OAuth 2LO/3LO token vault, workload identity, and least-privilege scope mismatches.
Fix AgentCore Memory retrieval returning empty or stale context by diagnosing the async long-term consolidation delay, short-term vs long-term memory, actor/session scoping, and namespace mismatches.
Fix AgentCore Runtime invocation timeouts by diagnosing cold-start init limits, the port 8080 and /invocations container contract, ARM64 builds, and 15-minute session timeouts.
Fix AgentCore Gateway 401/403 on MCP tool calls by checking the inbound JWT audience claim, the token itself, and the outbound OAuth2 credential provider.
Fix apply errors in a continuous AWS DMS task to Redshift or S3: diagnose stl_load_errors, VARCHAR overflow, IAM/S3 permissions, and instance sizing for an indefinitely-running CDC stream.
Fix rising CDC latency and AWS DMS task failures: diagnose CDCLatencySource vs CDCLatencyTarget, tables without a primary key, LOB truncation, and replication instance sizing.
Configure AWS RDS Proxy correctly: MaxConnectionsPercent, idle connection management, pinning avoidance, and the scenarios where RDS Proxy hurts more than it helps.
Fix a CloudWatch Alarm that won't trigger: INSUFFICIENT_DATA due to missing metrics, wrong evaluation periods, missing SNS permissions, and treat-missing-data configuration.
Fix Prometheus OOMKilled: identify high cardinality metrics, configure metric_relabel_configs to drop unnecessary series, set series limits, and right-size memory allocation.
Identify and reduce cloud sprawl: audit orphaned resources, unused accounts, untagged infrastructure, and duplicate services. CLI commands for AWS, GCP, and Azure with cost recovery estimates.
Diagnose and fix high CPU/memory usage by the Datadog Agent: disable unnecessary integrations, limit custom metrics, configure log collection, and switch to Cluster Agent.
Resolve GKE Binary Authorization deployment denials caused by missing attestations, wrong attestor key version, image digest mismatch, or missing break-glass annotation for emergency deploys.
Diagnose and fix API requests blocked by GCP VPC Service Controls perimeters: verify access levels, configure ingress/egress rules, analyze VPC-SC audit logs, and maintain PCI DSS compliance.
Diagnose and fix AWS Savings Plans utilization drop below 80%: identify causes (scale-down, workload migration, wrong region, seasonality) and strategies to recover commitment value.
Fixing vLLM throughput degradation after hours of operation: diagnosing KV cache fragmentation, enabling prefix caching, tuning gpu-memory-utilization, and scheduled restart.
Diagnosing a non-applying GCP CUD: verifying machine family, region, commitment type (compute vs memory-optimised), checking billing export, and fixing the issue.
Fix BigQuery queries stuck in queue: verify capacity commitment, assignment hierarchy, concurrency and slot autoscaling.
Resolving 429 errors in Azure OpenAI: identifying the bottleneck (TPM vs RPM), implementing retry logic, increasing quota, or switching to PTU.
Fixing Service Principal blocked by Conditional Access: Sign-in Logs analysis, excluding workload identities from policies, configuring Workload Identity Premium.
Fix Access Denied in Bedrock Knowledge Base sync: verify IAM role trust, KMS key policy, S3 bucket policy and VPC endpoint policy.
Fix Azure Database Migration Service connectivity errors: verify NSG rules, firewall configuration, Self-Hosted Integration Runtime and source network setup.
Fix DNS resolution in App Service with VNet Integration: configure Private DNS Zones, VNet links and WEBSITE_DNS_SERVER.
Fix Terraform drift after manual console changes: identify modified resources, import or reset to desired state, and implement drift detection with policies.
Fix ECS Task OOMKilled: identify the cause of memory limit breaches, right-size task definitions, configure memory reservations, and understand hard vs soft limits.
Fix vLLM inference pods crashing with CUDA OutOfMemoryError on EKS. Diagnose KV cache overflow, tune memory parameters, or switch to quantised model weights.
Fix OpenSearch Serverless vector search latency degradation caused by OCU starvation. Diagnose saturation, tune k-NN parameters, and scale OCUs.
Fix Bedrock InvokeModel failures with ThrottlingException or ModelStreamErrorException by diagnosing quota limits, implementing backoff, and requesting increases.
Diagnosing and fixing Terraform state drift: reconciling manual changes, terraform import, refresh-only plan. Safe force-unlock of a state locked in DynamoDB.
AWS environment security audit procedure, from IAM review, through network configuration, to encryption and monitoring. A ready-to-use checklist for recurring assessments.
Resolve GKE node pool stuck in PROVISIONING state caused by quota exhaustion, insufficient pod IP ranges, or machine type unavailability in the target zone.
Deploy Cloudflare Tunnel (cloudflared) as a Kubernetes Deployment to expose internal ClusterIP services without opening inbound ports. Works identically on EKS, GKE, AKS, and on-prem.
Diagnose and fix ALB 502 Bad Gateway errors caused by target group health check failures - security groups, health check paths, timeouts, and deploy draining.
Complete migration guide from deprecated aad-pod-identity to AKS Workload Identity Federation using OIDC issuer, managed identities, and federated credentials.
Fix a stuck Terraform state lock when terraform plan or apply fails with Error acquiring the state lock after a crashed or timed-out operation.
Fix ECS Fargate tasks that keep stopping with Essential container exited, OutOfMemoryError, or CannotPullContainerError by identifying the root cause in stopped task metadata and CloudWatch logs.
Fix RDS PostgreSQL connection limit errors by identifying idle connections, terminating leaked sessions, and implementing connection pooling with RDS Proxy or PgBouncer.
Fix the 10 most common high-risk items (HRI) from AWS Well-Architected Reviews. Step by step: IAM root access, encryption, backups, Multi-AZ, monitoring, logging - with CLI commands and Terraform.
Fixing 504 Timeout on API Gateway caused by Lambda cold start: diagnosing VPC ENI, optimising package size, provisioned concurrency, SnapStart for Java.
Quick reference mapping AWS container services (ECS, Fargate, ECR, App Runner) to their Azure equivalents (Container Apps, ACI, ACR, Web App for Containers) with guidance on when each applies.
Run a production Supabase stack using plain Docker Compose without the CLI. Full setup with custom registries, environment variables, and inter-service networking.
Fix asymmetric routing and traffic hairpinning when running multiple AWS Site-to-Site VPN connections in parallel to GCP Cloud VPN.
Fix supabase start failing with 403 Forbidden from public.ecr.aws by overriding the image registry to Docker Hub, GHCR, or your own private registry.
Fixing ThrottlingException 'too many tokens per day' in AWS Bedrock: analysing daily limits, retry with exponential backoff, requesting quota increase, fallback to alternative models.
Fixing 502 Bad Gateway on ALB with Lambda target: exceeding 1MB payload limit, ALB 29s timeout, malformed JSON response, Lambda cold start + ALB idle timeout.
Fix AWS Site-to-Site VPN throughput degradation caused by MTU fragmentation on IPsec tunnels connecting to GCP Cloud VPN.
Fixing networking issues when migrating from AWS Fargate to Azure Container Instances: DNS resolution, service discovery, NSG vs Security Groups, private endpoints.
Fix GCP Cross-Cloud Interconnect BGP session flapping caused by route advertisement limits, ASN conflicts, or hold timer mismatches.
Diagnose and fix AWS Site-to-Site VPN tunnel DOWN caused by IKE Phase 1/2 negotiation failure when connecting to GCP Cloud VPN.
Restore a BGP session over an Atman cross-connect when peering drops due to physical layer faults, MTU issues, or hold timer expiry.
Resolve PHP module conflicts in cPanel EasyApache4 when package dependencies clash during Apache rebuilds, causing build failures or broken PHP installations.
Rebuild an mdadm software RAID array on a Hetzner dedicated server after disk replacement using rescue mode, partitioning, and array reassembly.
Speed up Ceph recovery in Proxmox when an OSD is marked out and PG backfill crawls despite available disk bandwidth.
Fix the operation deadlock in ArgoCD when a stuck operation blocks all sync attempts with another operation is already in progress.
Fix the infinite OutOfSync loop in ArgoCD caused by the field ownership conflict between HPA and declarative spec.replicas management.
Diagnose and fix the etcdserver: request is too large error in ArgoCD when the Application CRD exceeds the 1 MB etcd object size limit.