<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Learn about DevOps, Linux, Containers, Kubernetes, CI/CD, AWS | SegFault]]></title><description><![CDATA[Learn about Devops, Security, Linux, Kubernetes,AWS, Terraform, Docker, and more. ]]></description><link>https://segfaultpw.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!1xLQ!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0301833-55a3-4a90-9c41-16084762ad60_512x512.png</url><title>Learn about DevOps, Linux, Containers, Kubernetes, CI/CD, AWS | SegFault</title><link>https://segfaultpw.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sat, 08 Aug 2026 23:18:21 GMT</lastBuildDate><atom:link href="https://segfaultpw.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Gabriel]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[segfaultpw@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[segfaultpw@substack.com]]></itunes:email><itunes:name><![CDATA[Gabriel]]></itunes:name></itunes:owner><itunes:author><![CDATA[Gabriel]]></itunes:author><googleplay:owner><![CDATA[segfaultpw@substack.com]]></googleplay:owner><googleplay:email><![CDATA[segfaultpw@substack.com]]></googleplay:email><googleplay:author><![CDATA[Gabriel]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[DevOps from Zero to Hero: Cost Optimization and What Comes Next]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-cost-optimization</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-cost-optimization</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!m0h-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!m0h-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!m0h-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!m0h-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!m0h-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!m0h-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!m0h-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/203015755?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!m0h-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!m0h-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!m0h-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!m0h-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F854c04ed-5288-4b88-b7c2-9fcd75333877_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://segfaultpw.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://segfaultpw.substack.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://segfaultpw.substack.com/p/devops-from-zero-to-hero-cost-optimization?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://segfaultpw.substack.com/p/devops-from-zero-to-hero-cost-optimization?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><h5><strong>Introduction</strong></h5><p>Welcome to article twenty, the final article of the DevOps from Zero to Hero series. Over the past nineteen articles we built an entire DevOps practice from scratch. We wrote a TypeScript API, learned version control, set up CI/CD pipelines, deployed to AWS, mastered Kubernetes, automated everything with GitOps, and added observability so we could actually see what was happening in production.</p><p>But there is one topic we have not covered yet, and it might be the one that gets you the most attention from leadership: cost. Cloud bills have a way of growing quietly in the background until someone notices a five-figure monthly invoice and starts asking hard questions. Cost optimization is not about being cheap. It is about spending intentionally and getting maximum value from every dollar.</p><p>In this article we will cover how to understand your AWS bill, identify common cost traps, right-size your resources, use Spot instances and Savings Plans, optimize Kubernetes costs, build a tagging strategy, set up cost monitoring, and manage dev/staging environments efficiently. Then we will wrap up the entire series with a full recap of everything we learned and talk about where to go from here.</p><p>Let's get into it.</p><h5><strong>Why cost matters: the rise of FinOps</strong></h5><p>When you are learning cloud in a personal account, costs feel manageable. A small EKS cluster, a few EC2 instances, and an RDS database might cost $100-300 per month. But in a real organization, those numbers multiply fast. Teams spin up resources and forget about them. Someone creates a NAT Gateway for testing and leaves it running for six months. A developer provisions an m5.4xlarge instance for a service that barely uses 10% of its CPU.</p><p>The cloud makes it incredibly easy to spend money. That is by design. There is no procurement process, no hardware to order, no six-week wait. You click a button and resources appear. This is powerful for speed, but dangerous for budgets.</p><p>This is where FinOps comes in. FinOps (Financial Operations) is a practice that brings financial accountability to cloud spending. It is not about cutting costs blindly. It is about making informed decisions about what to spend and why.</p><p>The core principles of FinOps are:</p><blockquote><ul><li><p><strong>Teams need to own their cloud costs</strong>: Just like DevOps made teams responsible for running their software, FinOps makes teams responsible for the cost of running it. If you deploy it, you should know what it costs.</p></li><li><p><strong>Decisions are driven by business value</strong>: Not every cost reduction is a good idea. Cutting your monitoring stack to save $500/month might cost you $50,000 when you miss an outage. Cost optimization is about value, not just spending less.</p></li><li><p><strong>Cloud is a variable cost model</strong>: Unlike on-premise where you buy servers and depreciate them over years, cloud costs change monthly. This means you need to review and optimize continuously, not just once a year.</p></li></ul></blockquote><p>Think of FinOps as the financial pillar of DevOps. You would not deploy code without testing it. You should not deploy infrastructure without understanding what it costs.</p><h5><strong>AWS Cost Explorer: understanding your bill</strong></h5><p>The first step in cost optimization is understanding where your money is going. AWS Cost Explorer is the primary tool for this. It is free and built into every AWS account.</p><p>To access it, go to the AWS Billing Console and click on Cost Explorer. The first time you enable it, it takes about 24 hours to populate historical data. After that, you get up to 12 months of spending history.</p><p>Here are the views you should use regularly:</p><p><strong>Monthly cost by service</strong></p><p>This is your starting point. Group by "Service" and set the time range to the last 3 months. You will immediately see which services are costing the most. In a typical Kubernetes-based setup, your top costs will usually be:</p><blockquote><ul><li><p><strong>EC2</strong> (including EKS worker nodes): Compute is almost always the biggest line item</p></li><li><p><strong>RDS</strong>: Database instances, especially if you run Multi-AZ</p></li><li><p><strong>NAT Gateway</strong>: Data transfer through NAT Gateways is surprisingly expensive</p></li><li><p><strong>EBS</strong>: Persistent volumes, snapshots, and unattached volumes</p></li><li><p><strong>S3</strong>: Storage and request costs</p></li><li><p><strong>Data Transfer</strong>: Cross-AZ and internet egress charges</p></li></ul></blockquote><p><strong>Cost by tag</strong></p><p>If you have a proper tagging strategy (we will cover this later), you can group costs by tag. This lets you answer questions like "How much does the staging environment cost?" or "What is team-alpha spending per month?" To use this view, you first need to activate your cost allocation tags in the Billing Console under Cost Allocation Tags.</p><p><strong>Daily cost trends</strong></p><p>Switch to daily granularity and look for spikes. A sudden jump in EC2 costs might mean someone launched a bunch of instances for a load test and forgot to terminate them. A spike in data transfer costs might indicate a misconfigured service that is pulling data across regions.</p><p>You can also use the AWS CLI to query cost data programmatically:</p><pre><code># Get last month's cost grouped by service
aws ce get-cost-and-usage \
  --time-period Start=2026-05-01,End=2026-06-01 \
  --granularity MONTHLY \
  --metrics "BlendedCost" \
  --group-by Type=DIMENSION,Key=SERVICE
</code></pre><pre><code># Get daily costs for the current month
aws ce get-cost-and-usage \
  --time-period Start=2026-06-01,End=2026-06-17 \
  --granularity DAILY \
  --metrics "BlendedCost"
</code></pre><h5><strong>Common cost traps</strong></h5><p>Every cloud environment has hidden costs waiting to surprise you. Here are the most common ones and how to find them.</p><p><strong>Forgotten resources</strong></p><p>These are resources that were created for a purpose but are no longer needed. They quietly accumulate charges every month.</p><blockquote><ul><li><p><strong>Unattached EBS volumes</strong>: When you terminate an EC2 instance, its EBS volumes might not be deleted automatically (depends on the DeleteOnTermination flag). These orphaned volumes cost money even when nothing is using them.</p></li><li><p><strong>Old EBS snapshots</strong>: Snapshots pile up over time. A daily snapshot policy on a 500GB volume creates 365 snapshots per year. At $0.05/GB-month, that adds up.</p></li><li><p><strong>Idle load balancers</strong>: A load balancer with no healthy targets still costs about $16-22/month. If you have abandoned ALBs from old projects, find them and delete them.</p></li><li><p><strong>NAT Gateways</strong>: Each NAT Gateway costs about $32/month just to exist, plus $0.045 per GB of data processed. If you have NAT Gateways in multiple AZs across multiple VPCs, that is hundreds of dollars per month doing nothing if those VPCs are inactive.</p></li><li><p><strong>Elastic IPs</strong>: An Elastic IP attached to a running instance is free. An Elastic IP not attached to anything costs $3.65/month. Small, but they add up.</p></li><li><p><strong>Unused ECR images</strong>: Container images in ECR cost $0.10/GB-month. If your CI pipeline pushes a new image on every commit and you never clean up old ones, storage costs grow linearly.</p></li></ul></blockquote><p>Find forgotten resources with these commands:</p><pre><code># Find unattached EBS volumes
aws ec2 describe-volumes \
  --filters Name=status,Values=available \
  --query 'Volumes[*].{ID:VolumeId,Size:Size,Created:CreateTime}' \
  --output table

# Find Elastic IPs not associated with anything
aws ec2 describe-addresses \
  --query 'Addresses[?AssociationId==`null`].{IP:PublicIp,AllocID:AllocationId}' \
  --output table

# Find load balancers with no targets
aws elbv2 describe-target-groups \
  --query 'TargetGroups[*].{ARN:TargetGroupArn,Name:TargetGroupName}' \
  --output table
</code></pre><p><strong>Oversized instances</strong></p><p>This is the most common cost trap. Teams pick an instance type when they first deploy a service and never revisit it. That m5.xlarge you chose "just in case" might be running at 5% CPU utilization. You could be on a t3.medium and save 75%.</p><p><strong>Idle dev/staging environments</strong></p><p>Your staging environment runs 24/7 but your team works 8 hours a day, 5 days a week. That means staging is idle 76% of the time. If staging costs $2,000/month, you are wasting about $1,500/month on compute that nobody is using.</p><p><strong>Cross-AZ data transfer</strong></p><p>Data transfer between Availability Zones costs $0.01/GB in each direction ($0.02/GB round trip). This sounds tiny, but a chatty microservice architecture with services spread across AZs can generate terabytes of cross-AZ traffic. This is often the most surprising line item on an AWS bill.</p><h5><strong>Right-sizing: matching resources to actual usage</strong></h5><p>Right-sizing means adjusting your compute resources to match what your workload actually needs. It is the highest-impact cost optimization you can do because compute is usually your biggest expense.</p><p><strong>Step 1: Gather metrics</strong></p><p>Before you can right-size anything, you need data. Use CloudWatch to understand your actual resource utilization:</p><pre><code># Get average CPU utilization for an instance over the last 7 days
aws cloudwatch get-metric-statistics \
  --namespace AWS/EC2 \
  --metric-name CPUUtilization \
  --dimensions Name=InstanceId,Value=i-0abc123def456789 \
  --start-time 2026-06-10T00:00:00Z \
  --end-time 2026-06-17T00:00:00Z \
  --period 3600 \
  --statistics Average Maximum \
  --output table
</code></pre><p>Look at both the average and the maximum. If your average CPU is 10% and your max is 25%, you have significant room to downsize. If your average is 10% but your max spikes to 95%, you might need that capacity for peak loads (or you might need to investigate what causes those spikes).</p><p><strong>Step 2: Use AWS Compute Optimizer</strong></p><p>AWS Compute Optimizer analyzes your CloudWatch metrics and recommends instance types that would better fit your workload. Enable it in the AWS Console under Compute Optimizer. It is free for basic recommendations.</p><p>It will tell you things like: "This m5.xlarge instance averages 8% CPU utilization. A t3.medium would save 75% while still providing sufficient capacity." These recommendations are a great starting point, but always validate them against your application's actual requirements. Memory-intensive applications might need more RAM than CPU, for example.</p><p><strong>Step 3: Right-size gradually</strong></p><p>Do not downsize everything at once. Pick your most over-provisioned instances, downsize them one at a time, and monitor for a week. If performance is fine, move to the next one. If you see issues, scale back up. Right-sizing is iterative, not a one-time event.</p><pre><code># Change instance type (requires stop/start)
aws ec2 stop-instances --instance-ids i-0abc123def456789
aws ec2 modify-instance-attribute \
  --instance-id i-0abc123def456789 \
  --instance-type '{"Value":"t3.medium"}'
aws ec2 start-instances --instance-ids i-0abc123def456789
</code></pre><p>For EKS worker nodes managed by a node group, you would update the launch template or node group configuration instead:</p><pre><code># Update managed node group instance type
aws eks update-nodegroup-config \
  --cluster-name my-cluster \
  --nodegroup-name my-nodegroup \
  --scaling-config minSize=2,maxSize=6,desiredSize=3
</code></pre><h5><strong>Spot instances and Karpenter</strong></h5><p>Spot instances let you use unused EC2 capacity at up to 90% discount compared to on-demand prices. The trade-off is that AWS can reclaim them with a 2-minute warning when it needs the capacity back. This sounds scary, but with the right architecture, Spot is one of the most effective cost optimization strategies available.</p><p><strong>How Spot works</strong></p><p>When AWS has unused capacity in a particular instance type and AZ, it makes that capacity available as Spot instances at a reduced price. The price fluctuates based on supply and demand but is typically 60-90% cheaper than on-demand. When AWS needs that capacity back (a "Spot interruption"), your instance gets a 2-minute warning and then is terminated.</p><p><strong>When to use Spot</strong></p><blockquote><ul><li><p><strong>Stateless workloads</strong>: Web servers, API servers, and workers that do not store data locally are perfect for Spot. If an instance gets interrupted, the load balancer routes traffic to other instances.</p></li><li><p><strong>Batch processing</strong>: Jobs that can be checkpointed and restarted work well on Spot.</p></li><li><p><strong>CI/CD runners</strong>: Build agents are short-lived by nature and can tolerate interruptions.</p></li><li><p><strong>Development and staging environments</strong>: These do not need the same reliability guarantees as production.</p></li></ul></blockquote><p><strong>When NOT to use Spot</strong></p><blockquote><ul><li><p><strong>Databases</strong>: Losing a database instance mid-transaction is a bad day.</p></li><li><p><strong>Stateful workloads without replication</strong>: If losing an instance means losing data, do not put it on Spot.</p></li><li><p><strong>Single-instance workloads</strong>: If you only have one instance and it gets interrupted, your service is down.</p></li></ul></blockquote><p><strong>Mixing on-demand and Spot</strong></p><p>The best practice is to run a baseline of on-demand instances that can handle your minimum expected load, and use Spot for everything above that. For example, if your API needs at least 3 instances to handle normal traffic but scales to 10 during peak hours, run 3 on-demand and let the remaining 7 be Spot.</p><p><strong>Karpenter for Kubernetes</strong></p><p>If you are running EKS, Karpenter is the best way to use Spot instances with Kubernetes. Karpenter is an open-source node provisioning tool that automatically selects the right instance types and purchase options (on-demand vs Spot) based on your pod requirements.</p><p>Here is a basic Karpenter NodePool configuration that mixes on-demand and Spot:</p><pre><code>apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: default
spec:
  template:
    spec:
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["on-demand", "spot"]
        - key: node.kubernetes.io/instance-type
          operator: In
          values:
            - m5.large
            - m5.xlarge
            - m5a.large
            - m5a.xlarge
            - m6i.large
            - m6i.xlarge
        - key: topology.kubernetes.io/zone
          operator: In
          values:
            - us-east-1a
            - us-east-1b
            - us-east-1c
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: default
  limits:
    cpu: "100"
    memory: 400Gi
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m
</code></pre><p>Karpenter will automatically diversify across multiple instance types and AZs to reduce the chance of simultaneous Spot interruptions. The <code>disruption</code> block tells Karpenter to consolidate underutilized nodes, which saves money by packing pods more efficiently.</p><p><strong>Handling Spot interruptions</strong></p><p>For graceful handling of Spot interruptions in Kubernetes, make sure your pods handle SIGTERM properly and have appropriate <code>terminationGracePeriodSeconds</code>. Karpenter integrates with the AWS Node Termination Handler to cordon and drain nodes before they are reclaimed.</p><h5><strong>Reserved Instances and Savings Plans</strong></h5><p>If you know you will need a certain amount of compute for the next 1-3 years, Reserved Instances (RIs) and Savings Plans offer significant discounts (up to 72%) in exchange for a commitment.</p><p><strong>Savings Plans vs Reserved Instances</strong></p><blockquote><ul><li><p><strong>Compute Savings Plans</strong>: You commit to a specific dollar amount of compute per hour (e.g., $10/hour) for 1 or 3 years. The discount applies across EC2, Fargate, and Lambda. This is the most flexible option.</p></li><li><p><strong>EC2 Instance Savings Plans</strong>: You commit to a specific instance family in a specific region (e.g., m5 in us-east-1). Higher discount than Compute Savings Plans but less flexible.</p></li><li><p><strong>Reserved Instances</strong>: You commit to a specific instance type, AZ, and tenancy. The highest discount but the least flexible. These are the legacy option and Savings Plans are generally recommended instead.</p></li></ul></blockquote><p><strong>When commitments make sense</strong></p><blockquote><ul><li><p><strong>Stable, predictable workloads</strong>: If your production database has been running on an r5.2xlarge for a year and will continue to do so, a Savings Plan is a no-brainer.</p></li><li><p><strong>Baseline compute</strong>: Commit to your minimum required compute. Use on-demand and Spot for anything above the baseline.</p></li><li><p><strong>After right-sizing</strong>: Always right-size first, then commit. There is nothing worse than committing to an oversized instance for 3 years.</p></li></ul></blockquote><p><strong>When to avoid commitments</strong></p><blockquote><ul><li><p><strong>New workloads</strong>: Wait until you understand the actual resource requirements (at least 2-3 months of data).</p></li><li><p><strong>Rapidly changing architectures</strong>: If you are migrating from EC2 to containers or from x86 to ARM, locking into commitments can backfire.</p></li><li><p><strong>Small amounts</strong>: The administrative overhead of managing RIs for a $50/month saving is not worth it.</p></li></ul></blockquote><p>A practical approach is to cover 60-70% of your steady-state compute with Savings Plans, handle the next 20% with on-demand, and use Spot for the remaining 10-20% that handles peak loads.</p><h5><strong>Kubernetes cost optimization</strong></h5><p>Kubernetes adds its own layer of cost complexity. Pods request resources, nodes provide them, and the gap between requested and actually used resources is wasted money.</p><p><strong>Resource requests and limits</strong></p><p>Every pod should have resource requests and limits defined. Requests tell the scheduler how much CPU and memory a pod needs. Limits cap how much it can use. The gap between what you request and what you actually use is waste.</p><pre><code>apiVersion: apps/v1
kind: Deployment
metadata:
  name: api
spec:
  replicas: 3
  template:
    spec:
      containers:
        - name: api
          image: my-api:latest
          resources:
            requests:
              cpu: "250m"
              memory: "256Mi"
            limits:
              cpu: "500m"
              memory: "512Mi"
</code></pre><p>The most common mistake is setting requests too high "just to be safe." If your API container uses 50m CPU on average but you request 500m, each pod wastes 450m of CPU. With 10 replicas, you are wasting 4.5 vCPUs, which could be an entire node worth of compute.</p><p>To find the right values, check actual usage with <code>kubectl top</code>:</p><pre><code># Check actual resource usage per pod
kubectl top pods -n my-namespace

# Check node-level resource utilization
kubectl top nodes

# Detailed resource allocation per node
kubectl describe node &lt;node-name&gt; | grep -A 5 "Allocated resources"
</code></pre><p>Set requests based on the P95 usage (what the pod actually uses 95% of the time) and limits at roughly 2x the request to handle bursts. Review and adjust these values every month.</p><p><strong>Namespace resource quotas</strong></p><p>Resource quotas prevent any single team or namespace from consuming more than its fair share of cluster resources. Without quotas, one team's runaway deployment can starve everyone else and force unnecessary cluster scaling.</p><pre><code>apiVersion: v1
kind: ResourceQuota
metadata:
  name: team-alpha-quota
  namespace: team-alpha
spec:
  hard:
    requests.cpu: "8"
    requests.memory: "16Gi"
    limits.cpu: "16"
    limits.memory: "32Gi"
    pods: "50"
    persistentvolumeclaims: "10"
</code></pre><p><strong>Cluster Autoscaler and Karpenter</strong></p><p>Both Cluster Autoscaler and Karpenter scale your node count based on pending pods, but they approach it differently:</p><blockquote><ul><li><p><strong>Cluster Autoscaler</strong>: Works with AWS Auto Scaling Groups. You predefine node group configurations (instance types, sizes). The autoscaler adds or removes nodes from these predefined groups. Simpler to set up but less flexible.</p></li><li><p><strong>Karpenter</strong>: Evaluates pending pods and provisions the optimal instance type on the fly. It can choose from a wide range of instance types and automatically bin-pack pods efficiently. More flexible and generally more cost-effective, but requires more initial configuration.</p></li></ul></blockquote><p>Whichever you use, make sure scale-down is enabled and tuned. By default, Cluster Autoscaler waits 10 minutes before removing an underutilized node. In a bursty environment, this delay means you are paying for idle nodes for 10 minutes after every traffic spike.</p><p><strong>Horizontal Pod Autoscaler (HPA)</strong></p><p>HPA scales your pod count based on metrics like CPU or custom metrics. This lets you run fewer pods during low-traffic periods and scale up during peaks, instead of running peak capacity 24/7.</p><pre><code>apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api
  minReplicas: 2
  maxReplicas: 10
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 60
</code></pre><h5><strong>Tagging strategy: tag everything</strong></h5><p>Tags are the foundation of cost visibility. Without tags, your AWS bill is one big number. With tags, you can answer "How much does each environment cost?", "Which team is spending the most?", and "What is the cost per customer?"</p><p><strong>Minimum required tags</strong></p><p>Every resource in your AWS account should have at least these tags:</p><blockquote><ul><li><p><strong>Environment</strong>: <code>production</code>, <code>staging</code>, <code>development</code></p></li><li><p><strong>Team</strong>: The team that owns the resource</p></li><li><p><strong>Service</strong>: The application or service name</p></li><li><p><strong>CostCenter</strong>: For chargeback or showback to business units</p></li><li><p><strong>ManagedBy</strong>: <code>terraform</code>, <code>manual</code>, <code>karpenter</code>, etc.</p></li></ul></blockquote><p><strong>Enforce tags with policies</strong></p><p>Tags only work if they are applied consistently. Use AWS Organizations tag policies or Terraform validation to enforce tagging:</p><pre><code># Terraform: enforce tags on all resources
variable "required_tags" {
  type = map(string)
  default = {
    Environment = ""
    Team        = ""
    Service     = ""
    ManagedBy   = "terraform"
  }
}

resource "aws_instance" "api" {
  ami           = "ami-0abc123def456789"
  instance_type = "t3.medium"

  tags = merge(var.required_tags, {
    Name        = "api-server"
    Environment = "production"
    Team        = "backend"
    Service     = "user-api"
  })
}
</code></pre><p>For a more robust approach, use an AWS Organizations tag policy:</p><pre><code>{
  "tags": {
    "Environment": {
      "tag_key": {
        "@@assign": "Environment"
      },
      "tag_value": {
        "@@assign": [
          "production",
          "staging",
          "development"
        ]
      },
      "enforced_for": {
        "@@assign": [
          "ec2:instance",
          "rds:db",
          "s3:bucket",
          "elasticloadbalancing:loadbalancer"
        ]
      }
    }
  }
}
</code></pre><p><strong>Activate cost allocation tags</strong></p><p>Creating tags is not enough. You also need to activate them as cost allocation tags in the Billing Console. Only activated tags appear in Cost Explorer for grouping and filtering. Go to Billing, then Cost Allocation Tags, find your tags, and click Activate. It takes up to 24 hours for activated tags to appear in Cost Explorer.</p><h5><strong>Cost monitoring: budgets and alerts</strong></h5><p>Setting up cost monitoring is like setting up application monitoring. You do not wait for users to report outages. You set up alerts. You should not wait for finance to report cost overruns either.</p><p><strong>AWS Budgets</strong></p><p>Create budgets for your total account spend and for each major service or environment:</p><pre><code># Create a monthly budget with email alerts
aws budgets create-budget \
  --account-id 123456789012 \
  --budget '{
    "BudgetName": "monthly-total",
    "BudgetLimit": {
      "Amount": "5000",
      "Unit": "USD"
    },
    "TimeUnit": "MONTHLY",
    "BudgetType": "COST"
  }' \
  --notifications-with-subscribers '[
    {
      "Notification": {
        "NotificationType": "ACTUAL",
        "ComparisonOperator": "GREATER_THAN",
        "Threshold": 80,
        "ThresholdType": "PERCENTAGE"
      },
      "Subscribers": [
        {
          "SubscriptionType": "EMAIL",
          "Address": "team@example.com"
        }
      ]
    },
    {
      "Notification": {
        "NotificationType": "FORECASTED",
        "ComparisonOperator": "GREATER_THAN",
        "Threshold": 100,
        "ThresholdType": "PERCENTAGE"
      },
      "Subscribers": [
        {
          "SubscriptionType": "EMAIL",
          "Address": "team@example.com"
        }
      ]
    }
  ]'
</code></pre><p>This creates a $5,000/month budget with two alerts: one when actual spend hits 80% of the budget, and another when the forecasted spend is projected to exceed the budget. The forecast alert is especially useful because it gives you time to act before you actually overspend.</p><p><strong>Weekly cost reviews</strong></p><p>Set up a weekly ritual where someone on the team reviews costs. It does not need to be a long meeting. A 15-minute check of Cost Explorer once a week is enough. Look for:</p><blockquote><ul><li><p><strong>Unexpected spikes</strong>: Anything that jumped significantly from the previous week</p></li><li><p><strong>New services</strong>: Any service that appeared in your bill that was not there before</p></li><li><p><strong>Trend lines</strong>: Is overall spending trending up? If so, is it proportional to growth?</p></li><li><p><strong>Idle resources</strong>: Any resources with zero or near-zero utilization</p></li></ul></blockquote><p>The person doing the review should rotate across the team. This builds cost awareness across the entire team, not just one designated cost watcher.</p><h5><strong>Dev/staging environment strategies</strong></h5><p>Development and staging environments are often the easiest place to cut costs because they do not need to be available 24/7 and they do not need production-grade resources.</p><p><strong>Scale down at night and on weekends</strong></p><p>If your team works 9am to 6pm on weekdays, your dev and staging environments are idle 73% of the time. Use scheduled scaling to shut them down outside working hours:</p><pre><code># Scale down EKS node group at night (run via cron or Lambda)
aws eks update-nodegroup-config \
  --cluster-name dev-cluster \
  --nodegroup-name dev-nodes \
  --scaling-config minSize=0,maxSize=3,desiredSize=0

# Scale up in the morning
aws eks update-nodegroup-config \
  --cluster-name dev-cluster \
  --nodegroup-name dev-nodes \
  --scaling-config minSize=1,maxSize=3,desiredSize=2
</code></pre><p>You can automate this with a Lambda function triggered by EventBridge on a schedule:</p><pre><code>{
  "schedule_expression": "cron(0 22 ? * MON-FRI *)",
  "description": "Scale down dev cluster at 10 PM",
  "action": "scale-down"
}
</code></pre><p><strong>Use smaller instances for non-production</strong></p><p>If production runs on m5.xlarge, staging can probably run on t3.medium. Dev can run on t3.small. The goal is not identical environments. It is environments that are similar enough to catch bugs but small enough to be affordable.</p><p><strong>Ephemeral environments</strong></p><p>Instead of running a persistent staging environment, consider spinning up short-lived environments for each pull request. The environment gets created when the PR is opened, runs integration tests, and gets destroyed when the PR is merged or closed. You only pay for the time someone is actively testing. Tools like Argo CD ApplicationSets or Terraform workspaces can automate this pattern.</p><p><strong>Single-node dev clusters</strong></p><p>For development, consider running a single-node Kubernetes cluster or using a local tool like kind or minikube. This avoids the EKS control plane cost ($73/month) and multi-node compute costs entirely for local development.</p><h5><strong>Putting it all together: a cost optimization checklist</strong></h5><p>Here is a practical checklist you can work through to optimize your cloud costs:</p><blockquote><ul><li><p><strong>Week 1</strong>: Enable Cost Explorer, activate cost allocation tags, create a basic budget with alerts</p></li><li><p><strong>Week 2</strong>: Audit for forgotten resources (unattached volumes, idle load balancers, unused Elastic IPs). Delete anything not needed</p></li><li><p><strong>Week 3</strong>: Analyze compute utilization with CloudWatch and Compute Optimizer. Identify right-sizing candidates</p></li><li><p><strong>Week 4</strong>: Right-size your most over-provisioned instances. Start with non-production</p></li><li><p><strong>Month 2</strong>: Implement tagging policies, set up scheduled scaling for dev/staging, evaluate Spot for stateless workloads</p></li><li><p><strong>Month 3</strong>: Review Kubernetes resource requests/limits, implement HPA, consider Karpenter. Evaluate Savings Plans for stable production workloads</p></li><li><p><strong>Ongoing</strong>: Weekly cost reviews, monthly optimization passes, quarterly Savings Plan evaluation</p></li></ul></blockquote><h5><strong>The complete series recap</strong></h5><p>We have covered a lot of ground in this series. Let's take a moment to look back at every article and what we learned in each one. If you missed any or want to revisit a topic, the links below will take you there.</p><blockquote><ul><li><p><strong>Article 1: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-what-it-actually-means">What It Actually Means</a></strong> - We started from the very beginning. What DevOps is, where it came from, the DORA metrics that measure it, and how DevOps relates to SRE and Platform Engineering.</p></li><li><p><strong>Article 2: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-your-first-typescript-api">Your First TypeScript API</a></strong> - We built a real application with Express and Docker. This gave us something concrete to deploy throughout the rest of the series.</p></li><li><p><strong>Article 3: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-version-control-for-teams">Version Control for Teams</a></strong> - We learned Git workflows, branching strategies, pull requests, and code review. The collaboration foundation for everything that followed.</p></li><li><p><strong>Article 4: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-automated-testing">Automated Testing</a></strong> - We wrote unit tests, integration tests, and learned the testing pyramid. No CI pipeline works without good tests.</p></li><li><p><strong>Article 5: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-your-first-ci-pipeline">Your First CI Pipeline</a></strong> - We set up GitHub Actions to automatically lint, test, and build our code on every push. Our first taste of automation.</p></li><li><p><strong>Article 6: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-aws-from-scratch">AWS from Scratch</a></strong> - We created an AWS account, set up IAM users and roles, understood regions and AZs, and got comfortable with the AWS CLI.</p></li><li><p><strong>Article 7: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-infrastructure-as-code">Infrastructure as Code with Terraform</a></strong> - We stopped clicking around in the console and started defining infrastructure as code. VPCs, subnets, security groups, all in Terraform.</p></li><li><p><strong>Article 8: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-deploying-to-ecs">Deploying to ECS with Fargate</a></strong> - We deployed our API to AWS for the first time using ECS and Fargate. Real cloud infrastructure running our real application.</p></li><li><p><strong>Article 9: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-secrets-and-config">Secrets and Config Management</a></strong> - We learned how to manage secrets safely with AWS Secrets Manager and SSM Parameter Store. No more hardcoded passwords.</p></li><li><p><strong>Article 10: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-dns-tls-and-networking">DNS, TLS, and Networking</a></strong> - We made our app reachable with a real domain, set up TLS certificates with ACM, and understood how networking ties everything together.</p></li><li><p><strong>Article 11: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-kubernetes-fundamentals">Kubernetes Fundamentals</a></strong> - We learned pods, deployments, services, and namespaces. The building blocks of container orchestration.</p></li><li><p><strong>Article 12: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-helm-charts">Helm Charts</a></strong> - We packaged our Kubernetes application with Helm, making it reusable and configurable across environments.</p></li><li><p><strong>Article 13: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-eks">EKS, Running Kubernetes on AWS</a></strong> - We set up a production-grade EKS cluster with Terraform, including managed node groups, IAM integration, and networking.</p></li><li><p><strong>Article 14: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-gitops-with-argocd">GitOps with ArgoCD</a></strong> - We implemented GitOps so that git became the single source of truth for our deployments. Push to git and ArgoCD handles the rest.</p></li><li><p><strong>Article 15: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-observability">Observability in Kubernetes</a></strong> - We set up Prometheus, Grafana, and structured logging. We learned about the three pillars: logs, metrics, and traces.</p></li><li><p><strong>Article 16: <a href="https://segfault.pw/blog/devops-from-zero-to-hero-the-complete-pipeline">CI/CD, The Complete Pipeline</a></strong> - We stitched everything together into a complete pipeline from pull request to production, with staging gates and manual approvals.</p></li><li><p><strong>Article 17: Security and Compliance</strong> - We covered container image scanning, RBAC policies, network policies, and how to bake security into every stage of the pipeline.</p></li><li><p><strong>Article 18: Disaster Recovery and High Availability</strong> - We learned multi-AZ deployments, backup strategies, RTO/RPO targets, and how to plan for the worst so your systems stay up.</p></li><li><p><strong>Article 19: Advanced Deployment Strategies</strong> - We explored canary deployments, blue/green deployments, feature flags, and progressive delivery patterns for zero-downtime releases.</p></li><li><p><strong>Article 20: Cost Optimization and What Comes Next (this article)</strong> - We learned how to understand, monitor, and optimize cloud costs, then wrapped up the entire series.</p></li></ul></blockquote><p>That is twenty articles, and if you followed along, you went from knowing nothing about DevOps to having a complete, production-grade pipeline with automated testing, infrastructure as code, Kubernetes, GitOps, observability, security, and cost optimization. That is a serious achievement.</p><h5><strong>What comes next</strong></h5><p>Finishing this series does not mean you are done learning. In many ways, you are just getting started. You now have a solid foundation, and there are several paths forward depending on your interests and career goals.</p><p><strong>Site Reliability Engineering (SRE)</strong></p><p>If you enjoyed the observability, monitoring, and reliability aspects of this series, SRE is a natural next step. SRE takes the DevOps principles we covered and adds rigorous engineering practices around reliability: SLIs, SLOs, error budgets, incident management, chaos engineering, and capacity planning.</p><p>We have an entire SRE series on this blog that picks up where this one leaves off. Start with <a href="https://segfault.pw/blog/sre-slis-slos-and-automations-that-actually-help">SRE: SLIs, SLOs, and Automations That Actually Help</a> and work through all fourteen articles.</p><p><strong>Platform Engineering</strong></p><p>If you found yourself thinking "I wish developers did not have to know all of this just to deploy their apps," Platform Engineering is for you. Platform teams build internal developer platforms that abstract away infrastructure complexity. You would build golden paths, self-service portals, and developer tooling that makes it easy for any developer to deploy, observe, and manage their applications without needing to understand every underlying component.</p><p><strong>Developer Experience (DX)</strong></p><p>Related to Platform Engineering, Developer Experience focuses on making developers productive and happy. Fast CI pipelines, great local development setups, clear documentation, easy onboarding. If you care about how people experience the tools and processes you build, DX is worth exploring.</p><p><strong>Certifications</strong></p><p>If you want to formalize your knowledge and signal your skills to employers, consider these certifications:</p><blockquote><ul><li><p><strong>AWS Solutions Architect Associate (SAA-C03)</strong>: Covers the core AWS services we used throughout this series. If you followed along and built everything, you already know about 70% of what is on this exam.</p></li><li><p><strong>Certified Kubernetes Administrator (CKA)</strong>: Validates your Kubernetes skills. The articles on Kubernetes fundamentals, Helm, and EKS gave you a strong head start.</p></li><li><p><strong>HashiCorp Terraform Associate</strong>: Covers the Terraform concepts we used for infrastructure as code. Probably the easiest of the three if you have been writing Terraform along with the series.</p></li></ul></blockquote><p>None of these certifications are required. Hands-on experience matters more than certificates. But they can be helpful for landing interviews, especially early in your career.</p><p><strong>Communities and resources</strong></p><p>Learning does not happen in isolation. Here are some communities and resources worth checking out:</p><blockquote><ul><li><p><strong>CNCF (Cloud Native Computing Foundation)</strong>: The organization behind Kubernetes, Prometheus, ArgoCD, and many other tools we used. Their landscape page gives you a map of the entire cloud native ecosystem.</p></li><li><p><strong>DevOps subreddits and forums</strong>: r/devops, r/kubernetes, and r/aws are active communities where people share experiences and help each other.</p></li><li><p><strong>KubeCon talks</strong>: The recorded talks from KubeCon are freely available on YouTube and cover everything from beginner to advanced topics.</p></li><li><p><strong>The SRE Book</strong>: Google's "Site Reliability Engineering" book is available free online at sre.google. It is the foundational text for SRE practices.</p></li><li><p><strong>"Accelerate" by Forsgren, Humble, and Kim</strong>: The book behind the DORA metrics. If you want to understand the research that proves DevOps practices work, this is the one.</p></li></ul></blockquote><h5><strong>Closing notes</strong></h5><p>This is the end of the DevOps from Zero to Hero series, and if you made it all the way here, I want to say something sincerely: well done. Twenty articles is a lot. Building all of this from scratch takes real commitment, and the fact that you stuck with it says a lot about you.</p><p>When we started this series, we talked about what DevOps actually means. Not the buzzword, not the job title, but the real idea: that the people who build software and the people who run it should work together, share responsibility, and use automation to move faster without sacrificing stability. Every article since then has been a practical expression of that idea. Automated tests, CI pipelines, infrastructure as code, Kubernetes, GitOps, observability, security, and now cost optimization. Each piece reinforces the others. Together, they form a complete practice.</p><p>But the most important thing you built is not a pipeline or a cluster. It is a way of thinking. You now approach problems differently. When you see a manual process, you think about automating it. When you see a deployment that requires SSH and prayer, you think about CI/CD. When someone says "it works on my machine," you think about containers. That mindset is more valuable than any specific tool, and it will serve you well no matter where your career takes you.</p><p>The cloud ecosystem will keep evolving. New tools will appear, some of what we covered will become outdated, and best practices will shift. That is fine. The fundamentals we covered (version control, testing, automation, infrastructure as code, observability, security, cost awareness) are timeless. The specific tools change, but the principles do not.</p><p>So go build something. Take what you learned here and apply it at work, on a side project, or in an open source contribution. The best way to solidify knowledge is to use it. And when you get stuck, remember that every expert you admire was once exactly where you are now.</p><p>Thank you for reading this series. I genuinely hope it helped you, and I hope you had as much fun following along as I had writing it. Until the next series!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Incident Response and On-Call]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-incident-response</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-incident-response</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Sun, 14 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!85md!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!85md!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!85md!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!85md!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!85md!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!85md!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!85md!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/202034430?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!85md!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!85md!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!85md!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!85md!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac05d85a-0e14-4a1e-8042-241dc1d8e384_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article nineteen of the DevOps from Zero to Hero series. In the previous articles we set up observability with Prometheus and Grafana, built dashboards, configured alerts, and deployed complete CI/CD pipelines. Everything is monitored and automated. But here is the question nobody wants to ask: what happens at 3am when an alert fires and your API is down?</p><p>That is what incident response is about. It is the human side of reliability. You can have the best monitoring in the world, but if nobody knows what to do when an alert goes off, it does not matter. In this article we are going to cover the fundamentals: what incidents are, how to classify them, how on-call rotations work, how to write runbooks that actually help, and how to learn from failures without blaming anyone.</p><p>This is a beginner-friendly introduction. If you want to go deeper into topics like incident commanders, SRE-specific practices, postmortems as code, and advanced on-call automation, check out the <a href="https://segfault.pw/blog/sre-incident-management-on-call-and-postmortems-as-code">SRE Incident Management</a> article from the SRE series. That article assumes you already understand the basics we cover here.</p><p>Let&#8217;s get into it.</p><h5><strong>What is an incident?</strong></h5><p>An incident is any unplanned event that disrupts or degrades a service for your users. Not every bug is an incident. A typo in a footer is a bug. Your payment processing being down for 500 users is an incident. The key distinction is user impact.</p><p>Most teams classify incidents by severity levels. The exact definitions vary between organizations, but here is a common framework:</p><blockquote><ul><li><p><strong>SEV1 (Critical)</strong>: Complete service outage or data loss. All or most users are affected. Example: the entire API is returning 500 errors, the database is unreachable, or customer data has been corrupted. This requires all hands on deck, immediately.</p></li><li><p><strong>SEV2 (Major)</strong>: Significant degradation but the service is partially working. Example: response times are 10x slower than normal, a key feature like checkout is broken, or 30% of requests are failing. This needs immediate attention from the on-call engineer.</p></li><li><p><strong>SEV3 (Minor)</strong>: A noticeable issue that affects a small number of users or a non-critical feature. Example: search suggestions are not loading, a dashboard widget shows stale data, or image uploads are slow. This should be addressed during business hours.</p></li><li><p><strong>SEV4 (Low)</strong>: A cosmetic issue or minor inconvenience with minimal user impact. Example: a tooltip has the wrong text, a non-critical background job is retrying more than usual, or a monitoring dashboard has a broken panel. This goes into the normal backlog.</p></li></ul></blockquote><p>The severity level determines everything else: who gets paged, how fast you need to respond, whether you need a status page update, and how much of the team gets pulled in. Getting this classification right is important because over-escalating burns people out and under-escalating lets problems grow.</p><h5><strong>The incident lifecycle</strong></h5><p>Every incident, regardless of severity, follows the same basic lifecycle. Understanding these phases helps you stay organized when things are stressful.</p><pre><code>  Detect &#9472;&#9472;&gt; Respond &#9472;&#9472;&gt; Mitigate &#9472;&#9472;&gt; Resolve &#9472;&#9472;&gt; Learn
    &#9474;           &#9474;            &#9474;            &#9474;           &#9474;
    &#9474;           &#9474;            &#9474;            &#9474;           &#9474;
  Alerts     Page the     Stop the     Fix the     Run a
  fire       on-call      bleeding     root        postmortem
             engineer                  cause</code></pre><p>Let&#8217;s walk through each phase:</p><p><strong>1. Detect</strong></p><p>Something tells you there is a problem. Ideally, your monitoring catches it before users do. In <a href="https://segfault.pw/blog/devops-from-zero-to-hero-observability">article fifteen</a> we set up Prometheus alerts that fire when error rates or latency exceed thresholds. Those alerts are your first line of detection. Other sources include health check failures, user reports, and automated smoke tests from your CI/CD pipeline.</p><p>The goal is simple: know about problems before your users tweet about them.</p><p><strong>2. Respond</strong></p><p>The alert reaches the on-call engineer through a tool like PagerDuty or OpsGenie. The on-call engineer acknowledges the alert (so the system knows someone is looking at it), assesses the severity, and decides if they need to pull in more people. For a SEV1, they might immediately start a war room. For a SEV3, they might just open a ticket and investigate during normal hours.</p><p><strong>3. Mitigate</strong></p><p>This is the most important phase and the one that trips up beginners. Mitigation is not about finding the root cause. It is about stopping the user impact as fast as possible. If your API is slow because a bad deployment went out, you roll back first and investigate later. If a database is overwhelmed, you scale it up or redirect traffic. Fix it enough to stop the pain, then figure out why it happened.</p><p>Common mitigation actions include:</p><blockquote><ul><li><p><strong>Rollback</strong>: Revert the last deployment if the issue started after a deploy</p></li><li><p><strong>Restart</strong>: Sometimes a simple pod restart clears a stuck process</p></li><li><p><strong>Scale up</strong>: Add more replicas or increase resource limits</p></li><li><p><strong>Feature flag</strong>: Disable a broken feature without rolling back everything</p></li><li><p><strong>Traffic shift</strong>: Route users to a healthy region or instance</p></li></ul></blockquote><p><strong>4. Resolve</strong></p><p>Once users are no longer affected, you can take the time to find and fix the actual root cause. Maybe the deployment was fine but it exposed a latent bug triggered by a specific data pattern. Maybe the database needs an index. Maybe the retry logic is creating a thundering herd. This is where you do the real engineering work.</p><p><strong>5. Learn</strong></p><p>After the incident is resolved, you run a postmortem. We will cover this in detail later in the article, but the short version is: you document what happened, build a timeline, identify the root cause, and create action items to prevent it from happening again. No blame. Just learning.</p><h5><strong>On-call basics</strong></h5><p>On-call means you are the designated person who responds when alerts fire outside of normal working hours (and often during them too). If you have never been on call before, the idea can be intimidating. Let&#8217;s break it down.</p><p><strong>What does being on call actually mean?</strong></p><p>When you are on call, you carry a phone (or have a laptop nearby) and you commit to responding to alerts within a defined time window, usually 5 to 15 minutes for critical alerts. You do not need to be sitting at your computer staring at dashboards. You can go to dinner, watch a movie, or sleep. But you need to be reachable and able to start investigating within the response time.</p><p><strong>Rotation schedules</strong></p><p>No one should be on call all the time. Teams set up rotations where the on-call responsibility passes from person to person on a regular schedule. Common patterns include:</p><blockquote><ul><li><p><strong>Weekly rotation</strong>: Person A is on call Monday to Monday, then Person B takes over. Simple and predictable. Works well for teams of 4 or more.</p></li><li><p><strong>Daily rotation</strong>: On-call shifts change every day. Less burden per shift but more handoffs. Good for teams that want to spread the load evenly.</p></li><li><p><strong>Follow-the-sun</strong>: If your team spans time zones, each region covers their daytime hours. Nobody gets woken up at 3am. This is the dream, but requires a globally distributed team.</p></li><li><p><strong>Primary and secondary</strong>: Two people are on call at the same time. The primary gets paged first. If they do not respond within the escalation window (say 10 minutes), the secondary gets paged. This provides a safety net.</p></li></ul></blockquote><p><strong>Compensation</strong></p><p>Being on call is work. Good organizations compensate for it. This can take different forms: extra pay for on-call shifts, time off after a busy on-call week, or a flat per-shift stipend. The specific approach varies, but the principle is clear: if you are asking someone to be available outside normal hours, you should recognize and compensate that time. Teams that do not compensate on-call eventually lose their best engineers.</p><p><strong>Handoff procedures</strong></p><p>When your on-call shift ends and someone else takes over, you should do a proper handoff. This means summarizing any ongoing issues, alerting quirks you noticed, or anything the next person should know. A quick message in Slack or a shared document works. The worst thing is inheriting an on-call shift with no context about what has been happening.</p><h5><strong>On-call tools</strong></h5><p>You need a tool that receives alerts from your monitoring system and routes them to the right person at the right time through the right channel (phone call, SMS, push notification, Slack). Here are the most common options:</p><blockquote><ul><li><p><strong>PagerDuty</strong>: The most established incident management platform. It handles alert routing, escalation policies, on-call schedules, and incident tracking. It integrates with everything: Prometheus, Grafana, AWS CloudWatch, Datadog, you name it. It is the industry standard but it is also the most expensive option.</p></li><li><p><strong>OpsGenie (by Atlassian)</strong>: Similar to PagerDuty in features, with strong integrations into the Atlassian ecosystem (Jira, Confluence, Statuspage). A solid choice if your team already uses Atlassian tools. Pricing is more accessible than PagerDuty.</p></li><li><p><strong>Grafana OnCall</strong>: An open-source option that integrates natively with Grafana. If you already use the Grafana stack for observability (as we set up in article fifteen), this is a natural fit. You can self-host it or use the Grafana Cloud managed version. It handles schedules, escalations, and routing, and it is free for self-hosted.</p></li></ul></blockquote><p>All three tools follow the same basic flow:</p><pre><code>Prometheus Alert
    &#9474;
    &#9660;
Alertmanager &#9472;&#9472;&gt; Webhook &#9472;&#9472;&gt; PagerDuty / OpsGenie / Grafana OnCall
                                &#9474;
                                &#9660;
                          On-call schedule
                                &#9474;
                                &#9660;
                         Page the on-call engineer
                         (phone, SMS, push, Slack)
                                &#9474;
                    &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9524;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
                    &#9474;                       &#9474;
              Acknowledged            Not acknowledged
              within SLA              within escalation window
                    &#9474;                       &#9474;
                    &#9660;                       &#9660;
              Engineer works          Page secondary /
              the incident           escalate to manager</code></pre><h5><strong>Setting up alerting to on-call tools</strong></h5><p>In <a href="https://segfault.pw/blog/devops-from-zero-to-hero-observability">article fifteen</a> we configured Prometheus alerts using PrometheusRule resources. Those alerts go to Alertmanager, which is part of the kube-prometheus-stack. Now we need to connect Alertmanager to an on-call tool so alerts actually reach a human.</p><p>Here is how you configure Alertmanager to send critical alerts to PagerDuty and non-critical alerts to a Slack channel:</p><pre><code># alertmanager-config.yaml
apiVersion: v1
kind: Secret
metadata:
  name: alertmanager-config
  namespace: monitoring
stringData:
  alertmanager.yaml: |
    global:
      resolve_timeout: 5m

    route:
      receiver: slack-default
      group_by: [alertname, namespace]
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h

      routes:
        # Critical alerts go to PagerDuty
        - receiver: pagerduty-critical
          match:
            severity: critical
          continue: false

        # Warning alerts go to Slack only
        - receiver: slack-warnings
          match:
            severity: warning
          continue: false

    receivers:
      - name: slack-default
        slack_configs:
          - api_url: "https://hooks.slack.com/services/YOUR/SLACK/WEBHOOK"
            channel: "#alerts"
            title: '{{ .GroupLabels.alertname }}'
            text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'

      - name: pagerduty-critical
        pagerduty_configs:
          - routing_key: "YOUR_PAGERDUTY_INTEGRATION_KEY"
            severity: critical
            description: '{{ .GroupLabels.alertname }}'
            details:
              namespace: '{{ .GroupLabels.namespace }}'
              summary: '{{ range .Alerts }}{{ .Annotations.summary }}{{ end }}'

      - name: slack-warnings
        slack_configs:
          - api_url: "https://hooks.slack.com/services/YOUR/SLACK/WEBHOOK"
            channel: "#alerts-low-priority"
            title: '{{ .GroupLabels.alertname }}'
            text: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'</code></pre><p>For OpsGenie, you would replace the <code>pagerduty_configs</code> with:</p><pre><code>      - name: opsgenie-critical
        opsgenie_configs:
          - api_key: "YOUR_OPSGENIE_API_KEY"
            message: '{{ .GroupLabels.alertname }}'
            priority: P1
            description: '{{ range .Alerts }}{{ .Annotations.description }}{{ end }}'</code></pre><p>For Grafana OnCall, you typically use a webhook receiver that points to your Grafana OnCall instance:</p><pre><code>      - name: grafana-oncall
        webhook_configs:
          - url: "https://oncall.your-grafana.com/integrations/v1/alertmanager/YOUR_ID/"
            send_resolved: true</code></pre><p>The key concept is routing. Not every alert should wake someone up. Route by severity: critical alerts page the on-call, warnings go to Slack, and informational alerts go to a low-priority channel. This is foundational for avoiding alert fatigue.</p><h5><strong>Alert fatigue</strong></h5><p>Alert fatigue is the number one killer of on-call programs. It happens when engineers get so many alerts that they start ignoring them. When you are getting paged 20 times a night, you stop taking alerts seriously. And when you stop taking alerts seriously, the one real outage that matters gets lost in the noise.</p><p>Here is the hard truth: too many alerts is worse than too few. With too few alerts, you might miss something, but at least when an alert fires, people pay attention. With too many alerts, people tune them out entirely, and you miss everything.</p><p><strong>Signs of alert fatigue:</strong></p><blockquote><ul><li><p><strong>High acknowledge rate, low action rate</strong>: People click &#8220;acknowledge&#8221; on alerts just to silence them, without actually investigating.</p></li><li><p><strong>Duplicate alerts</strong>: The same underlying issue triggers five different alerts, flooding the on-call with noise.</p></li><li><p><strong>Flapping alerts</strong>: An alert fires, resolves, fires, resolves, all within minutes. Each cycle generates a page.</p></li><li><p><strong>Low-value alerts</strong>: Alerts for things that do not require human action. &#8220;Disk usage at 70%&#8221; when you auto-scale at 80% is noise, not signal.</p></li><li><p><strong>After-hours pages for non-urgent issues</strong>: Getting woken up for a SEV4 that could wait until morning.</p></li></ul></blockquote><p><strong>How to fight alert fatigue:</strong></p><blockquote><ul><li><p><strong>Alert on symptoms, not causes</strong>: Page when users are affected (high error rate, slow responses), not when infrastructure metrics spike (CPU at 80%). We covered this in the observability article.</p></li><li><p><strong>Set meaningful thresholds</strong>: Do not set a latency alert at 200ms if your p99 is normally 180ms. Set it at a level that indicates a real problem, like 2x your normal p99.</p></li><li><p><strong>Use severity-based routing</strong>: Only page the on-call for critical alerts. Everything else goes to Slack or a ticket queue.</p></li><li><p><strong>Group related alerts</strong>: Configure Alertmanager&#8217;s <code>group_by</code> to combine related alerts into a single notification instead of five separate pages.</p></li><li><p><strong>Add inhibition rules</strong>: If the entire cluster is down, you do not need individual alerts for every service. An inhibition rule suppresses child alerts when a parent alert is firing.</p></li><li><p><strong>Review alerts regularly</strong>: Once a month, review all alerts that fired. Delete the ones that never led to action. Tune the thresholds on the ones that fire too often. This is ongoing maintenance, not a one-time task.</p></li></ul></blockquote><p>A good benchmark: the on-call engineer should get no more than two pages per on-call shift on average. If your team is consistently above that, you have a tuning problem, not a reliability problem.</p><h5><strong>Runbooks</strong></h5><p>A runbook is a documented procedure for handling a specific type of incident. When an alert fires and you are half-asleep at 3am, you do not want to figure out the debugging steps from scratch. You want a clear, step-by-step guide that tells you exactly what to check and what to do.</p><p><strong>What makes a good runbook:</strong></p><blockquote><ul><li><p><strong>It is linked from the alert</strong>: The alert annotation includes a URL to the runbook. One click from the page to the instructions.</p></li><li><p><strong>It starts with quick checks</strong>: The first steps should help you assess the severity and scope in under two minutes.</p></li><li><p><strong>It has concrete commands</strong>: Not &#8220;check the database&#8221; but &#8220;run this specific query and compare the result to this threshold.&#8221;</p></li><li><p><strong>It covers mitigation first, root cause second</strong>: Tell the engineer how to stop the bleeding before asking them to diagnose.</p></li><li><p><strong>It is kept up to date</strong>: A stale runbook is worse than no runbook because it gives false confidence. Review runbooks after every incident that uses them.</p></li></ul></blockquote><p>Here is a template you can use for any runbook:</p><pre><code># Runbook: [Alert Name]

## Overview
- **Alert**: [Name of the alert that links here]
- **Severity**: [SEV1/SEV2/SEV3]
- **Service**: [Which service is affected]
- **Last updated**: [Date]
- **Owner**: [Team or person responsible for this runbook]

## Quick assessment (do this first, under 2 minutes)
1. Check [dashboard link] for the current state
2. Run: `[specific command]` to confirm the issue
3. Determine scope: is it all users, a subset, or a single endpoint?

## Mitigation steps (stop the bleeding)
1. If this started after a recent deploy, rollback:
   `kubectl rollout undo deployment/[service] -n [namespace]`
2. If the issue is load-related, scale up:
   `kubectl scale deployment/[service] --replicas=[N] -n [namespace]`
3. [Any other quick fixes specific to this alert]

## Diagnosis (find the root cause)
1. Check logs: `kubectl logs -l app=[service] -n [namespace] --tail=100`
2. Check metrics: [specific PromQL query]
3. Check recent changes: [link to deploy history or git log]

## Escalation
- If you cannot mitigate within 30 minutes, escalate to [team/person]
- For data loss or security issues, immediately page [team/person]

## Previous incidents
- [Date]: [Brief description and link to postmortem]</code></pre><p>The most important part of this template is the &#8220;Quick assessment&#8221; section. It is what the on-call engineer reads first, bleary-eyed and trying to figure out if this is a real problem or a false alarm.</p><h5><strong>Practical example: API response time runbook</strong></h5><p>Let&#8217;s write a real runbook for one of the most common alerts: API response time exceeding 2 seconds. This connects directly to the Prometheus alerts we set up in article fifteen.</p><p>First, the alert rule that would trigger this:</p><pre><code># prometheus-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: api-latency-alerts
  namespace: monitoring
spec:
  groups:
    - name: api-latency
      rules:
        - alert: APIHighLatency
          expr: |
            histogram_quantile(0.99,
              sum(rate(http_request_duration_seconds_bucket{service="task-api"}[5m]))
              by (le)
            ) &gt; 2
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "API p99 latency is above 2 seconds"
            description: "The task-api p99 latency has been above 2s for 5 minutes."
            runbook_url: "https://wiki.example.com/runbooks/api-high-latency"</code></pre><p>Notice the <code>runbook_url</code> annotation. When this alert fires and reaches PagerDuty or OpsGenie, the runbook link is included in the notification. The on-call engineer can click it immediately.</p><p>Now the runbook itself:</p><pre><code># Runbook: APIHighLatency

## Overview
- **Alert**: APIHighLatency
- **Severity**: SEV2 (becomes SEV1 if latency exceeds 10s or error rate rises above 5%)
- **Service**: task-api
- **Last updated**: 2026-06-14
- **Owner**: Platform team

## Quick assessment (under 2 minutes)
1. Open the API dashboard: https://grafana.example.com/d/task-api
2. Check current p99 latency. Is it above 2s? How far above?
3. Check if error rate has also increased (indicates a deeper problem)
4. Check if the issue is isolated to one endpoint or all endpoints:
   Query: histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, route))

## Mitigation steps
### If latency started after a recent deploy:
1. Check the last deploy time:
   kubectl rollout history deployment/task-api -n production
2. If timing matches, roll back:
   kubectl rollout undo deployment/task-api -n production
3. Verify latency is recovering on the dashboard

### If latency is caused by high traffic:
1. Check current replica count and CPU usage:
   kubectl top pods -l app=task-api -n production
2. Scale up if replicas are at resource limits:
   kubectl scale deployment/task-api --replicas=6 -n production
3. Verify new pods are healthy:
   kubectl get pods -l app=task-api -n production

### If latency is caused by slow database queries:
1. Check database connection pool usage:
   Query: pg_stat_activity_count{datname="taskapi"}
2. Check for long-running queries:
   SELECT pid, now() - pg_stat_activity.query_start AS duration, query
   FROM pg_stat_activity
   WHERE state != 'idle' ORDER BY duration DESC LIMIT 10;
3. If a single query is blocking, consider cancelling it:
   SELECT pg_cancel_backend(&lt;pid&gt;);

## Diagnosis
1. Check logs for slow request patterns:
   kubectl logs -l app=task-api -n production --tail=200 | grep -i "slow\|timeout"
2. Check traces in Jaeger/Tempo for high-latency requests
3. Compare with normal baseline: typical p99 is 200-400ms
4. Review recent PRs merged to main for query changes or new endpoints

## Escalation
- If not mitigated within 30 minutes: page the backend team lead
- If database-related: page the DBA or infrastructure team
- If latency exceeds 10s or error rate &gt; 5%: escalate to SEV1

## Previous incidents
- 2026-05-20: Slow queries after migration added missing index. Fixed with CREATE INDEX.
- 2026-04-15: Memory leak caused GC pauses. Fixed with Node.js version upgrade.</code></pre><p>This runbook is specific, actionable, and organized by likelihood. The on-call engineer does not need to guess. They follow the steps, check the relevant data, and take action based on what they find.</p><h5><strong>Communication during incidents</strong></h5><p>When something is broken, people want to know. Your users, your support team, your leadership. Good communication during incidents reduces panic, builds trust, and lets you focus on fixing the problem instead of answering &#8220;is it fixed yet?&#8221; messages from twelve different people.</p><p><strong>Status pages</strong></p><p>A public (or internal) status page is the single source of truth during an incident. Tools like Statuspage (by Atlassian), Cachet (open source), or even a simple static page give your users a place to check instead of flooding your support channels.</p><p>A good status update includes:</p><blockquote><ul><li><p><strong>What is affected</strong>: &#8220;The checkout API is experiencing slow response times&#8221;</p></li><li><p><strong>Current status</strong>: &#8220;Investigating / Identified / Monitoring / Resolved&#8221;</p></li><li><p><strong>Impact</strong>: &#8220;Some users may experience delays when completing purchases&#8221;</p></li><li><p><strong>Next update</strong>: &#8220;We will provide an update in 30 minutes or when we have more information&#8221;</p></li></ul></blockquote><p>Keep updates factual and concise. Do not speculate about root causes in public updates. Say &#8220;we have identified the issue and are implementing a fix&#8221; rather than &#8220;we think the database index got corrupted.&#8221;</p><p><strong>Internal communication</strong></p><p>For your team and stakeholders, you need more detail. Most teams use a dedicated Slack channel per incident (for example, <code>#inc-2026-06-14-api-latency</code>). This keeps the conversation focused and creates a written record you can reference in the postmortem.</p><p>In the incident channel, post regular updates even if there is nothing new. &#8220;Still investigating, no new findings&#8221; is better than silence. Silence makes people nervous.</p><p><strong>War rooms</strong></p><p>For SEV1 incidents, teams often open a video call (sometimes called a war room or bridge call) where everyone working on the incident can communicate in real time. The key rules for war rooms:</p><blockquote><ul><li><p><strong>Keep it focused</strong>: Only people actively working on the incident should be in the call. Observers can follow the Slack channel.</p></li><li><p><strong>Designate a communication lead</strong>: One person handles all external updates so the engineers can focus on fixing things.</p></li><li><p><strong>Document decisions</strong>: Someone should be writing down what is happening, what has been tried, and what the current plan is. This becomes your postmortem timeline.</p></li></ul></blockquote><h5><strong>Blameless postmortems</strong></h5><p>A postmortem (also called a retrospective or incident review) is a structured analysis of what happened during an incident. The word &#8220;blameless&#8221; is the most important part. A blameless postmortem focuses on systems and processes, not on individuals.</p><p><strong>Why blameless matters</strong></p><p>If people are afraid they will be punished for causing an incident, they will hide information, avoid taking risks, and not report near-misses. Blame creates a culture of fear and silence. Blamelessness creates a culture of transparency and learning. The person who made the change that caused the outage is often the person who best understands the system and can help prevent it from happening again. You want them talking openly, not defending themselves.</p><p>This does not mean ignoring accountability. It means recognizing that most incidents are caused by system flaws (bad tooling, missing guardrails, unclear processes), not by people being careless.</p><p><strong>Postmortem template</strong></p><pre><code># Incident Postmortem: [Title]

## Summary
- **Date**: [When the incident occurred]
- **Duration**: [How long it lasted]
- **Severity**: [SEV level]
- **Impact**: [Who was affected and how]
- **Authors**: [Who wrote this postmortem]

## Timeline (all times in UTC)
- 14:32 - Monitoring alert fires: APIHighLatency
- 14:35 - On-call engineer acknowledges the alert
- 14:38 - Engineer checks dashboard, confirms p99 latency at 4.2s
- 14:42 - Identifies that latency spike started at 14:25, correlating with deploy abc123
- 14:45 - Initiates rollback of deployment
- 14:48 - Rollback complete, latency begins recovering
- 14:55 - Latency back to normal (p99 at 280ms)
- 14:58 - Incident marked as resolved

## Root cause
A database migration in commit abc123 added a new column to the orders table without an index.
The /orders endpoint performs a filter query on this column, which caused a full table scan on
every request. Under normal traffic, this increased p99 latency from 250ms to over 4 seconds.

## What went well
- Alert fired within 7 minutes of the deploy
- On-call engineer responded within 3 minutes
- Rollback was fast and effective (under 5 minutes from decision to recovery)
- Status page was updated within 10 minutes

## What went wrong
- The migration did not include an index for the new column
- No load testing was done against the staging database with realistic data volumes
- The staging database has 1,000 rows; production has 2 million, so the performance difference
  was not visible in staging

## Action items
- [ ] Add an index to the new column (owner: backend team, due: 2026-06-16)
- [ ] Add a CI check that flags migrations without indexes on queried columns (owner: platform team, due: 2026-06-30)
- [ ] Seed staging database with realistic data volumes (owner: platform team, due: 2026-07-15)
- [ ] Add a latency check to the post-deploy smoke tests (owner: platform team, due: 2026-06-30)</code></pre><p>Notice the structure. The timeline is factual and precise. The root cause is technical, not personal. &#8220;What went well&#8221; is just as important as &#8220;what went wrong&#8221; because it reinforces the things your team should keep doing. And the action items are specific, assigned, and have deadlines.</p><p><strong>Running the postmortem meeting</strong></p><p>Schedule the postmortem within 48 hours of the incident while memories are fresh. Keep it to 30-60 minutes. The facilitator (usually not someone directly involved in the incident) walks through the timeline and asks questions:</p><blockquote><ul><li><p>&#8220;What information did you have at this point?&#8221;</p></li><li><p>&#8220;What did you try and why?&#8221;</p></li><li><p>&#8220;What would have helped you resolve this faster?&#8221;</p></li><li><p>&#8220;Were there signals we missed that could have caught this earlier?&#8221;</p></li></ul></blockquote><p>The goal is to understand the system, not to judge decisions made under pressure. People make reasonable decisions based on the information they have at the time. If the system made it easy to deploy a migration without an index, the fix is a better system, not a lecture.</p><h5><strong>Building a healthy on-call culture</strong></h5><p>On-call does not have to be miserable. I have seen teams where on-call is dreaded and teams where it is manageable and even rewarding. The difference comes down to culture and investment.</p><p><strong>Reasonable expectations</strong></p><blockquote><ul><li><p><strong>Frequency</strong>: Nobody should be on call more than one week in four. If your team is too small for that rotation, you need to hire, share the rotation with another team, or reduce your on-call scope.</p></li><li><p><strong>Workload</strong>: The on-call engineer should be able to do their regular work during calm on-call shifts. If on-call is so busy that they cannot write code during the day, your alerts need tuning.</p></li><li><p><strong>Sleep</strong>: Getting paged once a night is acceptable occasionally. Getting paged three or four times every night is a systemic problem. Track after-hours pages as a metric and set a goal to reduce them.</p></li></ul></blockquote><p><strong>Practice incidents</strong></p><p>The worst time to learn incident response is during a real incident. Practice with game days or tabletop exercises. A game day is a planned exercise where you intentionally break something (in a controlled way) and practice the response. A tabletop exercise is where you walk through an incident scenario verbally without actually breaking anything.</p><pre><code>Example tabletop scenario:

"It is 2am on a Tuesday. You get paged for APIHighLatency.
 You check the dashboard and see p99 latency at 8 seconds.
 Error rate is at 12%. The last deploy was 6 hours ago.

 What do you do first?
 What do you check?
 Who do you contact?
 How do you communicate with stakeholders?"</code></pre><p>These exercises build muscle memory. When a real incident happens, the on-call engineer is not thinking &#8220;what do I do?&#8221; They are thinking &#8220;I have done this before, let me follow the process.&#8221;</p><p><strong>Handoff quality</strong></p><p>A good handoff between on-call shifts includes:</p><blockquote><ul><li><p><strong>Active incidents</strong>: Anything still ongoing or recently resolved</p></li><li><p><strong>Recent alerts</strong>: Alerts that fired and were handled, with context</p></li><li><p><strong>Known issues</strong>: Things that might page you but are already being worked on</p></li><li><p><strong>Environment changes</strong>: Recent deployments, infrastructure changes, or maintenance windows</p></li></ul></blockquote><p>A quick 15-minute call or a structured Slack message at handoff time prevents a lot of confusion.</p><p><strong>Investing in tooling</strong></p><p>Every time someone gets paged for something that could have been automated, that is a failure of tooling. Track your incidents and look for patterns. If the same issue keeps happening and the runbook is always &#8220;restart the pod,&#8221; automate the restart. If a particular alert always turns out to be a false positive, fix the alert. On-call should be for problems that genuinely need a human brain, not for tasks a script could handle.</p><h5><strong>Advanced topics</strong></h5><p>We have covered the fundamentals of incident response in this article, but there is much more to explore as your team and systems grow:</p><blockquote><ul><li><p><strong>Incident commander role</strong>: For SEV1 incidents, a dedicated incident commander coordinates the response, manages communication, and makes decisions about escalation. This role is separate from the engineers doing the technical work.</p></li><li><p><strong>SRE practices</strong>: Error budgets, SLO-based alerting, and toil reduction are advanced concepts that build on everything we covered here.</p></li><li><p><strong>Postmortems as code</strong>: Version-controlled postmortem templates, automated timeline generation, and action item tracking integrated into your project management tool.</p></li><li><p><strong>Chaos engineering</strong>: Intentionally injecting failures to test your incident response process before real incidents happen.</p></li></ul></blockquote><p>For a deep dive into all of these topics, check out the <a href="https://segfault.pw/blog/sre-incident-management-on-call-and-postmortems-as-code">SRE Incident Management</a> article. It covers incident commander workflows, on-call automation with Kubernetes operators, postmortem templates as code managed through GitOps, and advanced alerting strategies.</p><h5><strong>Closing notes</strong></h5><p>Incident response is not just about tools and processes. It is about people. It is about making sure the person who gets paged at 3am has what they need to solve the problem: clear alerts, good runbooks, the right access, and the confidence that comes from practice.</p><p>In this article we covered what incidents are and how to classify them with severity levels, the five phases of the incident lifecycle, how on-call rotations work and how to make them fair, setting up alerting from Prometheus to PagerDuty or OpsGenie, why alert fatigue is dangerous and how to fight it, how to write runbooks that actually help, communication best practices during incidents, blameless postmortems that drive improvement, and building a culture where on-call is sustainable.</p><p>The most important takeaway is this: invest in your incident response process before you need it. Write the runbooks, tune the alerts, practice the scenarios, and run the postmortems. When the real incident happens, you will be ready.</p><p>In the next and final article of the series, we will bring everything together and look at what comes after mastering the fundamentals.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Security Hardening]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-security-hardening</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-security-hardening</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BD0_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BD0_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BD0_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!BD0_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!BD0_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!BD0_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BD0_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/202034432?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BD0_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!BD0_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!BD0_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!BD0_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F25523c45-64bf-42a2-92d3-2b5fb1954a2c_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://segfaultpw.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://segfaultpw.substack.com/subscribe?"><span>Subscribe now</span></a></p><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article eighteen of the DevOps from Zero to Hero series. Over the past seventeen articles we have built an application, tested it, containerized it, deployed it to Kubernetes, set up GitOps with ArgoCD, added observability, and assembled a complete CI/CD pipeline. Everything works. But there is a question we have been skirting around the whole time: is any of this secure?</p><p>Security is not a feature you bolt on at the end. It is a practice you weave into every layer of your pipeline, your infrastructure, and your daily habits. The good news is that you do not need to be a security expert to get the basics right. Most real-world breaches come from simple mistakes: leaked credentials, unpatched dependencies, containers running as root, overly permissive access. These are all preventable with a checklist and some automation.</p><p>This article is intentionally a beginner-friendly checklist, not a deep dive. If you want comprehensive coverage of topics like OPA Gatekeeper, Falco, or policy-as-code frameworks, check out the <a href="https://segfault.pw/blog/sre-security-as-code">SRE Security as Code</a> article. For a thorough walk-through of Kubernetes RBAC at the API level, see the <a href="https://segfault.pw/blog/rbac-deep-dive">RBAC Deep Dive</a>. Here we are going to focus on the practical things every project should do from day one.</p><p>Let&#8217;s get into it.</p><h5><strong>The shift-left security mindset</strong></h5><p>The term &#8220;shift left&#8221; means moving security checks earlier in the development lifecycle. Instead of discovering a vulnerability in production (or worse, after a breach), you catch it during development or in your CI pipeline. The earlier you catch a problem, the cheaper and faster it is to fix.</p><p>Think of it like this. If you find a bug while writing code, it takes you five minutes to fix. If you find it in code review, it takes thirty minutes because you have to context-switch. If you find it in staging, it takes hours because now QA is involved. If you find it in production, it takes days and might involve an incident. Security issues follow the same curve, except the stakes are higher because a security issue can expose your users&#8217; data.</p><p>Shifting left does not mean you stop doing security reviews in production. It means you add automated checks at every stage so that the obvious stuff never makes it that far. Your CI pipeline becomes your first line of defense.</p><p>The pipeline stages where security checks belong:</p><blockquote><ul><li><p><strong>Code time</strong>: Linters, IDE plugins, pre-commit hooks that catch hardcoded secrets or insecure patterns before you even push</p></li><li><p><strong>Pull request</strong>: SAST tools, dependency scanners, and secret detection run as CI checks on every PR</p></li><li><p><strong>Build time</strong>: Container image scanning, SBOM generation, base image verification</p></li><li><p><strong>Deploy time</strong>: Kubernetes admission controllers, Pod Security Standards, RBAC enforcement</p></li><li><p><strong>Runtime</strong>: Network policies, audit logging, runtime threat detection (covered in the SRE series)</p></li></ul></blockquote><p>The rest of this article walks through each of these stages with practical examples you can add to your project today.</p><h5><strong>SAST: Static Application Security Testing</strong></h5><p>SAST tools analyze your source code without running it. They look for patterns that are known to cause security issues: SQL injection, cross-site scripting (XSS), command injection, insecure cryptography, hardcoded credentials, and more.</p><p>The key thing to understand is that SAST does not find every bug. It finds common patterns that match known vulnerability signatures. Think of it as a spell checker for security. It catches the obvious mistakes so you can focus your manual review time on the subtle ones.</p><p><strong>Semgrep</strong> is one of the best tools for this. It is open source, supports many languages, and has a huge library of community rules. You can also write your own rules for patterns specific to your codebase.</p><p>Here is how to add Semgrep to your GitHub Actions pipeline:</p><pre><code># .github/workflows/security.yml
name: Security Checks

on:
  pull_request:
    branches: [main]
  push:
    branches: [main]

jobs:
  sast:
    name: SAST Scan
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Run Semgrep
        uses: semgrep/semgrep-action@v1
        with:
          config: &gt;-
            p/security-audit
            p/secrets
            p/owasp-top-ten
        env:
          SEMGREP_APP_TOKEN: ${{ secrets.SEMGREP_APP_TOKEN }}</code></pre><p>The <code>p/security-audit</code>, <code>p/secrets</code>, and <code>p/owasp-top-ten</code> are rule packs that cover the most common vulnerability patterns. Semgrep will scan your code and report any matches as comments on your pull request.</p><p>For JavaScript and TypeScript projects, you should also add ESLint security plugins:</p><pre><code>npm install --save-dev eslint-plugin-security eslint-plugin-no-secrets</code></pre><pre><code>{
  "plugins": ["security", "no-secrets"],
  "extends": ["plugin:security/recommended"],
  "rules": {
    "no-secrets/no-secrets": "error",
    "security/detect-eval-with-expression": "error",
    "security/detect-non-literal-fs-filename": "warn",
    "security/detect-possible-timing-attacks": "warn"
  }
}</code></pre><p>These plugins run as part of your normal linting step, so they catch issues before code even gets to the PR stage.</p><h5><strong>Dependency scanning</strong></h5><p>Your application code is probably 10% of the code that actually runs. The other 90% comes from dependencies. And those dependencies have their own dependencies (transitive dependencies). A vulnerability in a deeply nested transitive dependency can be just as dangerous as one in your own code.</p><p>This is not theoretical. The Log4Shell vulnerability (CVE-2021-44228) was in a logging library that was a transitive dependency in thousands of Java applications. Most teams did not even know they were using it until the CVE dropped.</p><p><strong>npm audit</strong> is the simplest starting point for Node.js projects:</p><pre><code># Check for known vulnerabilities
npm audit

# Fix automatically where possible
npm audit fix

# Fail CI if there are high or critical vulnerabilities
npm audit --audit-level=high</code></pre><p>Add this to your CI pipeline:</p><pre><code>  dependency-scan:
    name: Dependency Scan
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Install dependencies
        run: npm ci

      - name: Run npm audit
        run: npm audit --audit-level=high</code></pre><p><strong>Dependabot</strong> is built into GitHub and automatically creates pull requests when new vulnerability patches are available. Enable it by adding a configuration file:</p><pre><code># .github/dependabot.yml
version: 2
updates:
  - package-ecosystem: "npm"
    directory: "/"
    schedule:
      interval: "weekly"
    open-pull-requests-limit: 10
    reviewers:
      - "your-team"

  - package-ecosystem: "docker"
    directory: "/"
    schedule:
      interval: "weekly"

  - package-ecosystem: "github-actions"
    directory: "/"
    schedule:
      interval: "weekly"</code></pre><p>Notice that we are scanning three ecosystems: npm packages, Docker base images, and GitHub Actions versions. Each one is a potential attack surface.</p><p><strong>Snyk</strong> is another popular option that provides deeper analysis than npm audit, including fix suggestions and prioritization based on exploitability. It has a free tier for open source projects.</p><p>The key habit here is: treat dependency updates as security maintenance, not optional chores. When Dependabot opens a PR, review it and merge it promptly. Stale dependencies are one of the most common attack vectors.</p><h5><strong>Container image scanning with Trivy</strong></h5><p>Your Docker images contain an entire operating system plus your application and its dependencies. Every package in that OS is a potential vulnerability. Trivy is an open-source scanner that checks your container images (and your Dockerfiles, and your Kubernetes manifests) for known vulnerabilities.</p><p>First, scan your Dockerfile for misconfigurations:</p><pre><code># Install Trivy
brew install trivy  # macOS
# or: sudo apt-get install trivy  # Ubuntu

# Scan a Dockerfile for misconfigurations
trivy config Dockerfile

# Scan a built image for vulnerabilities
trivy image myapp:latest

# Only show high and critical vulnerabilities
trivy image --severity HIGH,CRITICAL myapp:latest

# Fail if any critical vulnerabilities are found (useful for CI)
trivy image --severity CRITICAL --exit-code 1 myapp:latest</code></pre><p>Common issues Trivy catches in Dockerfiles:</p><blockquote><ul><li><p><strong>Running as root</strong>: Your container should use a non-root user. Add <code>USER nonroot</code> to your Dockerfile.</p></li><li><p><strong>Using latest tag</strong>: Always pin your base image to a specific version or digest.</p></li><li><p><strong>Missing health checks</strong>: Add a <code>HEALTHCHECK</code> instruction so orchestrators know when your app is unhealthy.</p></li><li><p><strong>Sensitive data in layers</strong>: Never <code>COPY</code> secrets into your image. Use build args or mount secrets at runtime.</p></li></ul></blockquote><p>Here is a CI job that scans your image after building it:</p><pre><code>  image-scan:
    name: Container Image Scan
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Build image
        run: docker build -t myapp:${{ github.sha }} .

      - name: Run Trivy scan
        uses: aquasecurity/trivy-action@0.28.0
        with:
          image-ref: myapp:${{ github.sha }}
          format: table
          exit-code: 1
          severity: HIGH,CRITICAL
          ignore-unfixed: true</code></pre><p>The <code>ignore-unfixed: true</code> flag skips vulnerabilities that do not have a fix available yet. This prevents your pipeline from blocking on issues you cannot actually resolve. You should still track unfixed vulnerabilities, but they should not break your build.</p><h5><strong>OIDC for CI/CD authentication</strong></h5><p>If your GitHub Actions workflows deploy to AWS (or any cloud provider), you need credentials. The old way was to store long-lived access keys as GitHub Secrets. The problem is that those keys never expire, they exist in multiple places, and if they leak, an attacker has persistent access to your AWS account.</p><p>OIDC (OpenID Connect) solves this by letting GitHub Actions request short-lived credentials directly from AWS. No long-lived keys stored anywhere. The credentials last for the duration of the workflow run and then they expire.</p><p>Here is how to set it up:</p><p><strong>Step 1: Create an OIDC identity provider in AWS</strong></p><pre><code># Create the OIDC provider (one-time setup)
aws iam create-open-id-connect-provider \
  --url https://token.actions.githubusercontent.com \
  --thumbprint-list "6938fd4d98bab03faadb97b34396831e3780aea1" \
  --client-id-list "sts.amazonaws.com"</code></pre><p><strong>Step 2: Create an IAM role with a trust policy</strong></p><pre><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Principal": {
        "Federated": "arn:aws:iam::ACCOUNT_ID::oidc-provider/token.actions.githubusercontent.com"
      },
      "Action": "sts:AssumeRoleWithWebIdentity",
      "Condition": {
        "StringEquals": {
          "token.actions.githubusercontent.com:aud": "sts.amazonaws.com"
        },
        "StringLike": {
          "token.actions.githubusercontent.com:sub": "repo:your-org/your-repo:ref:refs/heads/main"
        }
      }
    }
  ]
}</code></pre><p>The <code>Condition</code> block is important. It restricts which repository and branch can assume this role. Without it, any GitHub repository could use your AWS credentials.</p><p><strong>Step 3: Use OIDC in your workflow</strong></p><pre><code>jobs:
  deploy:
    runs-on: ubuntu-latest
    permissions:
      id-token: write   # Required for OIDC
      contents: read
    steps:
      - uses: actions/checkout@v4

      - name: Configure AWS credentials
        uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::ACCOUNT_ID::role/github-actions-deploy
          aws-region: us-east-1

      - name: Deploy
        run: |
          # These credentials are short-lived and scoped to this workflow run
          aws sts get-caller-identity
          # ... your deployment commands</code></pre><p>The <code>permissions.id-token: write</code> line is what enables OIDC. Without it, the workflow cannot request a token from GitHub&#8217;s OIDC provider.</p><p>This pattern works with AWS, GCP, Azure, and any cloud provider that supports OIDC. If your provider supports it, there is no reason to use long-lived keys.</p><h5><strong>Kubernetes RBAC basics</strong></h5><p>RBAC (Role-Based Access Control) controls who can do what in your Kubernetes cluster. The principle is simple: every user, service, and automation should have the minimum permissions it needs to do its job, and nothing more.</p><p>RBAC has four key resources:</p><blockquote><ul><li><p><strong>Role</strong>: Defines a set of permissions within a namespace. For example, &#8220;can read pods and services in the staging namespace.&#8221;</p></li><li><p><strong>ClusterRole</strong>: Same as Role but applies across the entire cluster. Use this for cluster-wide resources like nodes or namespaces.</p></li><li><p><strong>RoleBinding</strong>: Connects a Role to a user, group, or ServiceAccount within a namespace.</p></li><li><p><strong>ClusterRoleBinding</strong>: Connects a ClusterRole to a subject across the entire cluster.</p></li></ul></blockquote><p>Here is a basic example that gives a CI/CD ServiceAccount permission to manage deployments in a specific namespace:</p><pre><code># Create a ServiceAccount for your CI/CD pipeline
apiVersion: v1
kind: ServiceAccount
metadata:
  name: ci-deployer
  namespace: production
---
# Define what it can do
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: deployer-role
  namespace: production
rules:
  - apiGroups: ["apps"]
    resources: ["deployments"]
    verbs: ["get", "list", "update", "patch"]
  - apiGroups: [""]
    resources: ["pods"]
    verbs: ["get", "list", "watch"]
---
# Bind the role to the ServiceAccount
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: ci-deployer-binding
  namespace: production
subjects:
  - kind: ServiceAccount
    name: ci-deployer
    namespace: production
roleRef:
  kind: Role
  name: deployer-role
  apiGroup: rbac.authorization.k8s.io</code></pre><p>Notice that the Role only grants <code>get</code>, <code>list</code>, <code>update</code>, and <code>patch</code> on deployments. It does not grant <code>create</code> or <code>delete</code>. It also does not grant access to secrets, configmaps, or any other resource. This is the principle of least privilege in action.</p><p>Common mistakes to avoid:</p><blockquote><ul><li><p><strong>Using cluster-admin for everything</strong>: The <code>cluster-admin</code> ClusterRole gives full access to everything. Never bind it to service accounts used by applications or CI pipelines.</p></li><li><p><strong>Using default ServiceAccounts</strong>: Every namespace has a <code>default</code> ServiceAccount. If you do not create specific ones, all your pods share the same identity. Create dedicated ServiceAccounts for each application.</p></li><li><p><strong>Not auditing RBAC</strong>: Run <code>kubectl auth can-i --list --as=system:serviceaccount:production:ci-deployer</code> to verify what a ServiceAccount can actually do.</p></li></ul></blockquote><p>For a much deeper exploration of RBAC, including how it works at the HTTP API level with raw curl calls, check out the <a href="https://segfault.pw/blog/rbac-deep-dive">RBAC Deep Dive</a>.</p><h5><strong>Network Policies</strong></h5><p>By default, every pod in a Kubernetes cluster can talk to every other pod. This is convenient for development but terrible for security. If an attacker compromises one pod, they can reach every other service in the cluster.</p><p>Network Policies let you control which pods can communicate with which other pods. Think of them as firewall rules for your cluster&#8217;s internal network.</p><p><strong>Step 1: Start with a default deny policy</strong></p><p>This blocks all ingress traffic to pods in the namespace. Nothing can talk to anything unless you explicitly allow it.</p><pre><code>apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-ingress
  namespace: production
spec:
  podSelector: {}    # Applies to all pods in the namespace
  policyTypes:
    - Ingress</code></pre><p><strong>Step 2: Allow specific traffic paths</strong></p><p>Now you poke holes for the traffic that needs to flow. For example, let the API receive traffic from the ingress controller:</p><pre><code>apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-ingress-to-api
  namespace: production
spec:
  podSelector:
    matchLabels:
      app: api
  policyTypes:
    - Ingress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              name: ingress-system
        - podSelector:
            matchLabels:
              app: ingress-controller
      ports:
        - port: 3000
          protocol: TCP</code></pre><p>And let the API talk to the database:</p><pre><code>apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-api-to-db
  namespace: production
spec:
  podSelector:
    matchLabels:
      app: postgres
  policyTypes:
    - Ingress
  ingress:
    - from:
        - podSelector:
            matchLabels:
              app: api
      ports:
        - port: 5432
          protocol: TCP</code></pre><p>The pattern is always the same: deny everything by default, then allow only the specific paths your application needs. Document these paths. If you cannot explain why a network policy exists, it probably should not.</p><p>Important note: Network Policies require a CNI plugin that supports them. If you are using EKS, the default VPC CNI does not enforce Network Policies. You need to enable the Network Policy feature or use a CNI like Calico. Check your cluster&#8217;s CNI documentation.</p><h5><strong>Pod Security Standards</strong></h5><p>Pod Security Standards (PSS) define three profiles that control what pods are allowed to do at the security level:</p><blockquote><ul><li><p><strong>Privileged</strong>: No restrictions. Use this only for system-level pods like CNI plugins or storage drivers that genuinely need host access.</p></li><li><p><strong>Baseline</strong>: Prevents the most dangerous configurations like running as privileged, using host networking, or mounting the host filesystem. This is a reasonable default for most workloads.</p></li><li><p><strong>Restricted</strong>: The strictest profile. Requires running as non-root, drops all Linux capabilities, sets a read-only root filesystem, and more. This is what production applications should target.</p></li></ul></blockquote><p>The simplest way to enforce these is with namespace labels. Kubernetes has a built-in admission controller called Pod Security Admission that reads these labels and enforces the corresponding profile.</p><pre><code>apiVersion: v1
kind: Namespace
metadata:
  name: production
  labels:
    # Enforce: reject pods that violate the restricted profile
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/enforce-version: latest

    # Warn: log a warning for pods that violate restricted
    # (useful during migration to see what would break)
    pod-security.kubernetes.io/warn: restricted
    pod-security.kubernetes.io/warn-version: latest

    # Audit: add an audit annotation for baseline violations
    pod-security.kubernetes.io/audit: restricted
    pod-security.kubernetes.io/audit-version: latest</code></pre><p>When you apply the <code>enforce: restricted</code> label, Kubernetes will reject any pod that does not meet the restricted profile. For example, if your pod spec does not include <code>runAsNonRoot: true</code>, the pod will be rejected at admission time.</p><p>Here is what a pod spec looks like when it meets the restricted profile:</p><pre><code>apiVersion: v1
kind: Pod
metadata:
  name: secure-app
  namespace: production
spec:
  securityContext:
    runAsNonRoot: true
    runAsUser: 1000
    fsGroup: 1000
    seccompProfile:
      type: RuntimeDefault
  containers:
    - name: app
      image: myapp:1.0.0@sha256:abc123...
      securityContext:
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities:
          drop:
            - ALL
      resources:
        limits:
          memory: "128Mi"
          cpu: "500m"
        requests:
          memory: "64Mi"
          cpu: "250m"</code></pre><p>If you are migrating existing workloads, start with <code>warn</code> mode to see what would fail, fix the violations, and then switch to <code>enforce</code>. Do not jump straight to enforce on a production namespace unless you have tested every workload.</p><h5><strong>Secrets hygiene</strong></h5><p>Secrets management is covered in depth in the <a href="https://segfault.pw/blog/devops-from-zero-to-hero-secrets-and-config">Secrets and Config</a> article from this series. Here we are going to focus on the hygiene practices that prevent secrets from leaking in the first place.</p><p><strong>Never log secrets</strong>. This sounds obvious, but it happens all the time. A debug log statement prints the entire request object, which includes the Authorization header. A startup script echoes environment variables to verify configuration. An error handler dumps the full context, including database connection strings. All of these end up in your logging system, which is usually accessible to far more people than should have access to your secrets.</p><p>Practical rules:</p><blockquote><ul><li><p><strong>Redact sensitive fields in logging</strong>: Configure your logging library to redact fields like <code>password</code>, <code>token</code>, <code>secret</code>, <code>authorization</code>, and <code>cookie</code>. Most logging libraries support this.</p></li><li><p><strong>Never echo secrets in CI logs</strong>: If your CI pipeline needs a secret, use masked variables. GitHub Actions masks secrets automatically, but only if you reference them through <code>${{ secrets.NAME }}</code>. If you copy the value to a regular variable and echo it, the masking does not apply.</p></li><li><p><strong>Rotate secrets regularly</strong>: Set a rotation schedule. At minimum, rotate every 90 days. Rotate immediately if someone leaves the team or if you suspect a leak.</p></li><li><p><strong>Audit who has access</strong>: Periodically review who can read your secrets. In Kubernetes, check which ServiceAccounts have <code>get</code> or <code>list</code> on secrets. In GitHub, review who has access to repository secrets.</p></li><li><p><strong>Use short-lived tokens</strong>: Whenever possible, use tokens that expire. OIDC tokens, JWTs with short expiry, temporary AWS credentials. Long-lived tokens are a liability.</p></li></ul></blockquote><p>A quick way to check for hardcoded secrets in your codebase before they get committed:</p><pre><code># Install gitleaks
brew install gitleaks  # macOS

# Scan the current repo
gitleaks detect --source . --verbose

# Add as a pre-commit hook
# .pre-commit-config.yaml
repos:
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.18.0
    hooks:
      - id: gitleaks</code></pre><h5><strong>Supply chain security</strong></h5><p>Supply chain attacks target the tools and dependencies you use rather than your code directly. The SolarWinds attack, the Codecov breach, and the ua-parser-js npm hijack are all examples. You cannot eliminate supply chain risk entirely, but you can reduce your exposure significantly.</p><p><strong>Pin action versions by SHA, not tag</strong></p><p>GitHub Actions tags are mutable. A malicious actor who compromises a popular action&#8217;s repository can update the <code>v4</code> tag to point to malicious code, and every workflow using <code>actions/checkout@v4</code> would run it. Pinning by SHA makes your workflow reproducible and tamper-resistant:</p><pre><code># Instead of this (tag can be moved):
- uses: actions/checkout@v4

# Use this (SHA is immutable):
- uses: actions/checkout@b4ffde65f46336ab88eb53be808477a3936bae11  # v4.1.1</code></pre><p>Yes, it is less readable. Add a comment with the version number. Security is worth the trade-off.</p><p><strong>Verify base images</strong></p><p>Use official images from trusted registries. Pin them by digest, not just by tag:</p><pre><code># Instead of this (tag can be overwritten):
FROM node:20-alpine

# Use this (digest is content-addressable and immutable):
FROM node:20-alpine@sha256:abcdef1234567890...</code></pre><p>You can find the digest on Docker Hub or by running <code>docker inspect --format='{{index .RepoDigests 0}}' node:20-alpine</code>.</p><p><strong>SBOM (Software Bill of Materials)</strong></p><p>An SBOM is an inventory of every component in your application. When a new CVE drops, an SBOM tells you immediately whether you are affected. You do not have to go digging through <code>node_modules</code> or Docker layers.</p><p>Trivy can generate SBOMs:</p><pre><code># Generate an SBOM in SPDX format
trivy image --format spdx-json --output sbom.json myapp:latest

# Generate in CycloneDX format
trivy image --format cyclonedx --output sbom.xml myapp:latest</code></pre><p>Store your SBOM as a build artifact in your CI pipeline so you can reference it later when new vulnerabilities are disclosed.</p><p>For a more comprehensive treatment of supply chain security including Cosign image signing and Kyverno policies, see the <a href="https://segfault.pw/blog/sre-security-as-code">SRE Security as Code</a> article.</p><h5><strong>The practical security checklist</strong></h5><p>Here is a top-10 list of things every project should implement. These are ordered roughly by impact and effort, so start from the top and work your way down.</p><blockquote><ul><li><p><strong>1. Enable Dependabot or equivalent</strong>: Turn on automated dependency updates for all your ecosystems (npm, Docker, GitHub Actions). This takes five minutes and catches most known vulnerabilities automatically.</p></li><li><p><strong>2. Add secret scanning to your repo</strong>: Enable GitHub secret scanning or add gitleaks as a pre-commit hook. This prevents accidental credential leaks, which are the number one cause of breaches in small teams.</p></li><li><p><strong>3. Scan container images in CI</strong>: Add Trivy or a similar scanner to your build pipeline. Fail the build on critical vulnerabilities. This catches OS-level vulnerabilities that your language-level scanners miss.</p></li><li><p><strong>4. Use OIDC instead of long-lived keys</strong>: If your CI deploys to a cloud provider, switch to OIDC authentication. Remove any long-lived access keys from your GitHub Secrets.</p></li><li><p><strong>5. Run containers as non-root</strong>: Update your Dockerfiles to use a non-root user. Apply Pod Security Standards at the namespace level to enforce this cluster-wide.</p></li><li><p><strong>6. Implement network policies</strong>: Start with default-deny and explicitly allow the traffic your application needs. This limits blast radius if a pod gets compromised.</p></li><li><p><strong>7. Create dedicated RBAC roles</strong>: Stop using cluster-admin and default ServiceAccounts. Create specific Roles with minimum permissions for each workload and CI pipeline.</p></li><li><p><strong>8. Add SAST to your CI pipeline</strong>: Add Semgrep or equivalent SAST tooling. Even the default rule packs catch a surprising number of real issues.</p></li><li><p><strong>9. Pin your dependencies</strong>: Pin action versions by SHA, pin base images by digest, and use lock files for package managers. This protects against supply chain attacks.</p></li><li><p><strong>10. Rotate secrets on a schedule</strong>: Set calendar reminders to rotate API keys, database passwords, and service account tokens every 90 days. Automate rotation where possible.</p></li></ul></blockquote><p>You do not need to do all ten in one sprint. Start with items 1 through 4. They are quick wins with high impact. Then work through the rest as you mature your security posture.</p><h5><strong>Putting it all together in CI</strong></h5><p>Here is a complete security workflow that combines several of the checks we discussed. You can add this alongside your existing CI pipeline:</p><pre><code># .github/workflows/security.yml
name: Security

on:
  pull_request:
    branches: [main]
  push:
    branches: [main]
  schedule:
    # Run weekly even without code changes to catch new CVEs
    - cron: "0 8 * * 1"

jobs:
  sast:
    name: SAST
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@b4ffde65f46336ab88eb53be808477a3936bae11  # v4.1.1

      - name: Semgrep
        uses: semgrep/semgrep-action@v1
        with:
          config: &gt;-
            p/security-audit
            p/secrets
            p/owasp-top-ten

  dependencies:
    name: Dependency Scan
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@b4ffde65f46336ab88eb53be808477a3936bae11  # v4.1.1

      - name: Install dependencies
        run: npm ci

      - name: npm audit
        run: npm audit --audit-level=high

  image-scan:
    name: Image Scan
    runs-on: ubuntu-latest
    needs: [sast, dependencies]  # Only scan if code checks pass
    steps:
      - uses: actions/checkout@b4ffde65f46336ab88eb53be808477a3936bae11  # v4.1.1

      - name: Build image
        run: docker build -t myapp:${{ github.sha }} .

      - name: Trivy scan
        uses: aquasecurity/trivy-action@0.28.0
        with:
          image-ref: myapp:${{ github.sha }}
          format: table
          exit-code: 1
          severity: HIGH,CRITICAL
          ignore-unfixed: true

      - name: Generate SBOM
        if: github.ref == 'refs/heads/main'
        run: |
          trivy image --format spdx-json \
            --output sbom.json myapp:${{ github.sha }}

      - name: Upload SBOM
        if: github.ref == 'refs/heads/main'
        uses: actions/upload-artifact@v4
        with:
          name: sbom
          path: sbom.json</code></pre><p>A few things to notice:</p><blockquote><ul><li><p><strong>Scheduled runs</strong>: The <code>cron</code> trigger runs the scan weekly even if no code changes. New CVEs are published constantly, and a dependency that was clean last week might have a critical vulnerability today.</p></li><li><p><strong>Actions pinned by SHA</strong>: We practice what we preach. The checkout action is pinned to a specific commit.</p></li><li><p><strong>SBOM as artifact</strong>: On main branch builds, we generate and store an SBOM so we have a record of exactly what went into each release.</p></li><li><p><strong>Fail fast</strong>: The image scan only runs if SAST and dependency checks pass. No point scanning an image if the code itself has issues.</p></li></ul></blockquote><h5><strong>Closing notes</strong></h5><p>Security is not a project with a finish date. It is a practice, like testing or code review. The goal is not to make your system impenetrable (nothing is), but to make it hard enough that attackers move on to easier targets, and to limit the damage when something does get through.</p><p>In this article we covered the shift-left security mindset, SAST with Semgrep and ESLint plugins, dependency scanning with npm audit and Dependabot, container image scanning with Trivy, OIDC authentication for CI/CD, Kubernetes RBAC basics, network policies with default deny, Pod Security Standards and namespace enforcement, secrets hygiene practices, supply chain security with pinned versions and SBOMs, and a practical top-10 security checklist.</p><p>Every topic here was covered at the checklist level. If you want to go deeper, the <a href="https://segfault.pw/blog/sre-security-as-code">SRE Security as Code</a> article covers OPA Gatekeeper, Falco runtime security, Cosign image signing, and policy-as-code frameworks. The <a href="https://segfault.pw/blog/rbac-deep-dive">RBAC Deep Dive</a> article walks through RBAC at the Kubernetes API level with raw HTTP calls.</p><p>Start with the checklist. Pick the top three or four items that your project is missing and implement them this week. Security is one of those things where doing something is infinitely better than doing nothing.</p><p>In the next article we will cover disaster recovery and backup strategies, the final layer of protection when everything else fails.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Database Migrations and Zero-Downtime Deployments]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-database-migrations</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-database-migrations</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BJeX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BJeX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BJeX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!BJeX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!BJeX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!BJeX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BJeX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201090743?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BJeX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!BJeX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!BJeX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!BJeX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74f38831-3ec8-44f7-ad83-d2502bc07e8d_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article seventeen of the DevOps from Zero to Hero series. In the previous articles we built a complete CI/CD pipeline, set up observability, and deployed our TypeScript API to Kubernetes. Everything works, the pipeline is green, and deploys are smooth. But there is one topic we have been quietly avoiding: the database.</p><p>Deploying new application code is relatively straightforward. You build a new image, roll it out, and if something goes wrong, you roll back. But database changes are different. They are stateful. You cannot just &#8220;undo&#8221; a column drop. They affect every instance of your application simultaneously. And if you get the ordering wrong, you can take down your entire service.</p><p>In this article we will cover what database migrations are and how they work, how to write safe migrations using Prisma, the expand-contract pattern for making backwards-compatible schema changes, zero-downtime deployment strategies in Kubernetes, health checks and readiness probes, and rollback strategies for when things go wrong. By the end, you will know how to ship database changes confidently, even under production traffic.</p><p>Let&#8217;s get into it.</p><h5><strong>Why database changes are the riskiest part of deployments</strong></h5><p>Application code is stateless. If you deploy a bad version, you roll back to the previous container image and the problem is gone. The old code runs exactly as it did before. But databases are stateful. Once you drop a column, that data is gone. Once you rename a table, every query that references the old name breaks instantly.</p><p>Here is what makes database changes so dangerous:</p><blockquote><ul><li><p><strong>They are shared state</strong>: Every pod, every instance, every replica reads from the same database. A schema change affects all of them at once. You cannot do a gradual rollout of a database change the way you can with application code.</p></li><li><p><strong>They are hard to reverse</strong>: Adding a column is easy to undo (just drop it). But dropping a column, changing a column type, or deleting data? Those operations are destructive. You cannot &#8220;undelete&#8221; a column and get the data back.</p></li><li><p><strong>Ordering matters</strong>: If your application code expects a column that does not exist yet, it crashes. If your migration removes a column that old application code still references, it crashes. The sequencing between code deploys and schema changes is critical.</p></li><li><p><strong>They hold locks</strong>: Many schema changes (especially on large tables) acquire locks that block reads or writes. A migration that takes 30 seconds to run on your dev database might take 30 minutes on a production table with millions of rows, locking out all traffic.</p></li></ul></blockquote><p>The core challenge is this: during a deployment, you will have old code and new code running at the same time. Your database schema must be compatible with both versions simultaneously. This constraint drives every decision we will make in this article.</p><h5><strong>Migration fundamentals</strong></h5><p>A database migration is a versioned, incremental change to your database schema. Instead of manually running SQL statements against your database, you write migration files that describe the change, and a migration tool applies them in order.</p><p>Every migration has two parts:</p><blockquote><ul><li><p><strong>Up</strong>: The forward change. Create a table, add a column, create an index. This is what runs when you apply the migration.</p></li><li><p><strong>Down</strong>: The reverse change. Drop the table, remove the column, remove the index. This is what runs when you roll back the migration.</p></li></ul></blockquote><p>Migration files are typically named with a timestamp or sequence number so the tool knows what order to run them in:</p><pre><code>migrations/
  20260601120000_create_users_table/
    migration.sql
  20260602090000_add_email_to_orders/
    migration.sql
  20260603140000_create_audit_log/
    migration.sql</code></pre><p>The migration tool keeps track of which migrations have been applied in a special table (usually called <code>_prisma_migrations</code> or <code>schema_migrations</code>). When you run <code>migrate</code>, it checks which migrations are pending and applies them in order. This gives you a complete, auditable history of every schema change, the ability to reproduce your database schema from scratch, and a mechanism to roll back changes when needed.</p><h5><strong>Setting up migrations with Prisma</strong></h5><p>We will use Prisma as our TypeScript ORM and migration tool. Prisma takes a different approach from traditional migration tools: you define your schema in a declarative file, and Prisma generates the SQL migrations for you.</p><p>First, install Prisma in your project:</p><pre><code>npm install prisma --save-dev
npm install @prisma/client

# Initialize Prisma with PostgreSQL
npx prisma init --datasource-provider postgresql</code></pre><p>This creates a <code>prisma/schema.prisma</code> file. Let&#8217;s define our first model:</p><pre><code>// prisma/schema.prisma
generator client {
  provider = "prisma-client-js"
}

datasource db {
  provider = "postgresql"
  url      = env("DATABASE_URL")
}

model User {
  id        Int      @id @default(autoincrement())
  name      String
  email     String   @unique
  createdAt DateTime @default(now())
  updatedAt DateTime @updatedAt
  orders    Order[]
}

model Order {
  id        Int      @id @default(autoincrement())
  amount    Float
  status    String   @default("pending")
  userId    Int
  user      User     @relation(fields: [userId], references: [id])
  createdAt DateTime @default(now())
}</code></pre><p>Now generate the first migration:</p><pre><code>npx prisma migrate dev --name create_users_and_orders</code></pre><p>Prisma compares your schema file to the current database state, generates the SQL, and applies it. The generated migration looks like this:</p><pre><code>-- CreateTable
CREATE TABLE "User" (
    "id" SERIAL NOT NULL,
    "name" TEXT NOT NULL,
    "email" TEXT NOT NULL,
    "createdAt" TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP,
    "updatedAt" TIMESTAMP(3) NOT NULL,

    CONSTRAINT "User_pkey" PRIMARY KEY ("id")
);

-- CreateTable
CREATE TABLE "Order" (
    "id" SERIAL NOT NULL,
    "amount" DOUBLE PRECISION NOT NULL,
    "status" TEXT NOT NULL DEFAULT 'pending',
    "userId" INTEGER NOT NULL,
    "createdAt" TIMESTAMP(3) NOT NULL DEFAULT CURRENT_TIMESTAMP,

    CONSTRAINT "Order_pkey" PRIMARY KEY ("id")
);

-- CreateIndex
CREATE UNIQUE INDEX "User_email_key" ON "User"("email");

-- AddForeignKey
ALTER TABLE "Order" ADD CONSTRAINT "Order_userId_fkey"
  FOREIGN KEY ("userId") REFERENCES "User"("id") ON DELETE RESTRICT ON UPDATE CASCADE;</code></pre><p>Now let&#8217;s add a column. Say we need a <code>phone</code> field on the <code>User</code> model:</p><pre><code>model User {
  id        Int      @id @default(autoincrement())
  name      String
  email     String   @unique
  phone     String?  // nullable, so existing rows are fine
  createdAt DateTime @default(now())
  updatedAt DateTime @updatedAt
  orders    Order[]
}</code></pre><pre><code>npx prisma migrate dev --name add_phone_to_users</code></pre><p>The generated SQL:</p><pre><code>-- AlterTable
ALTER TABLE "User" ADD COLUMN "phone" TEXT;</code></pre><p>Notice that the column is nullable (<code>TEXT</code> without <code>NOT NULL</code>). This is important. If we made it required, the migration would fail because existing rows would not have a value for the new column. Making new columns nullable (or giving them a default value) is one of the most basic safe migration patterns.</p><p>Now let&#8217;s rename a column. Say we want to rename <code>name</code> to <code>fullName</code>. In Prisma, you use the <code>@map</code> attribute to rename the database column without changing the Prisma field name:</p><pre><code>model User {
  id        Int      @id @default(autoincrement())
  fullName  String   @map("full_name")
  email     String   @unique
  phone     String?
  createdAt DateTime @default(now())
  updatedAt DateTime @updatedAt
  orders    Order[]

  @@map("users")
}</code></pre><p>But hold on. If we just rename the column directly, any application code that still queries the old column name will break. This is exactly the kind of dangerous operation we need to handle with the expand-contract pattern, which we will cover next.</p><h5><strong>Safe migration patterns</strong></h5><p>The golden rule of safe migrations is: never make a breaking change in a single deploy. Instead, break it into multiple steps where each step is backwards-compatible.</p><p>Here are the patterns you should follow:</p><p><strong>Adding a column (safe)</strong></p><p>Adding a nullable column or a column with a default value is always safe. Old code ignores the new column. New code can use it.</p><pre><code>-- Safe: nullable column, old code ignores it
ALTER TABLE "User" ADD COLUMN "phone" TEXT;

-- Safe: column with default, old code ignores it
ALTER TABLE "Order" ADD COLUMN "currency" TEXT NOT NULL DEFAULT 'USD';</code></pre><p><strong>Adding an index (safe, but watch lock time)</strong></p><p>Creating an index is safe from a compatibility standpoint, but on large tables it can lock the table for a long time. Use <code>CONCURRENTLY</code> in PostgreSQL to avoid blocking writes:</p><pre><code>-- This locks the table until the index is built (dangerous on large tables)
CREATE INDEX idx_orders_user_id ON "Order" ("userId");

-- This builds the index without locking (safe for production)
CREATE INDEX CONCURRENTLY idx_orders_user_id ON "Order" ("userId");</code></pre><p><strong>Creating a new table (safe)</strong></p><p>New tables do not affect existing code at all. Always safe.</p><p><strong>Dropping a column (dangerous if done wrong)</strong></p><p>If you drop a column that old application code still reads from, those queries will fail. Never drop a column in the same deploy that still references it.</p><p><strong>Renaming a column (dangerous if done in one step)</strong></p><p>A column rename is essentially a drop-and-add from the application&#8217;s perspective. Old code queries the old name. New code queries the new name. During a rolling update, both versions run simultaneously, and one of them will always be broken.</p><h5><strong>The expand-contract pattern</strong></h5><p>The expand-contract pattern is the standard way to make breaking schema changes safely. It works in three phases:</p><p><strong>Phase 1: Expand (add the new thing)</strong></p><p>Add the new column alongside the old one. Update your application code to write to both columns. Deploy this change.</p><pre><code>-- Migration 1: Add the new column
ALTER TABLE "users" ADD COLUMN "full_name" TEXT;</code></pre><pre><code>// Application code writes to both columns
async function updateUser(id: number, name: string) {
  await prisma.user.update({
    where: { id },
    data: {
      name: name,      // old column (for old code still reading it)
      fullName: name,  // new column (for new code)
    },
  });
}</code></pre><p><strong>Phase 2: Migrate data</strong></p><p>Backfill the new column with data from the old column. This can be a migration script or a background job.</p><pre><code>-- Migration 2: Backfill existing data
UPDATE "users" SET "full_name" = "name" WHERE "full_name" IS NULL;</code></pre><p>At this point, both columns have the same data. Old code reads from <code>name</code>, new code reads from <code>full_name</code>, and everything works.</p><p><strong>Phase 3: Contract (remove the old thing)</strong></p><p>Once all application code has been updated to use the new column, and you have verified that no queries reference the old column, you can drop it.</p><pre><code>-- Migration 3: Drop the old column (only after all code uses full_name)
ALTER TABLE "users" DROP COLUMN "name";</code></pre><p>This three-phase approach means that at every point during the rollout, the database schema is compatible with both the old and new versions of your code. No downtime, no errors, no data loss.</p><p>Here is the timeline:</p><pre><code># Expand-contract timeline
#
# Deploy 1: Add "full_name" column, write to both columns
#   Old code: reads "name"       -&gt; works (column still exists)
#   New code: reads "full_name"  -&gt; works (column was just added)
#
# Deploy 2: Backfill "full_name" from "name"
#   All rows now have both columns populated
#
# Deploy 3: Remove all reads from "name", drop column
#   Old code: gone (fully rolled out)
#   New code: reads "full_name"  -&gt; works (only column left)</code></pre><p>Yes, this takes three deploys instead of one. That is the trade-off. Safety costs velocity, but it saves you from 3am incidents.</p><h5><strong>Dangerous patterns to avoid</strong></h5><p>Here are the schema changes that cause the most outages, and how to handle them instead:</p><blockquote><ul><li><p><strong>Renaming a column in a single deploy</strong>: This is a drop plus add. Use expand-contract instead. Add the new column, backfill, update code, then drop the old column.</p></li><li><p><strong>Changing a column type in place</strong>: Changing <code>VARCHAR(50)</code> to <code>TEXT</code> might seem harmless, but changing <code>TEXT</code> to <code>INTEGER</code> will fail if any rows contain non-numeric data. Add a new column with the new type, backfill with type conversion, switch code, then drop the old column.</p></li><li><p><strong>Adding a NOT NULL constraint without a default</strong>: If you add <code>NOT NULL</code> to an existing column that has null values, the migration will fail. First backfill all nulls, then add the constraint.</p></li><li><p><strong>Dropping a table that is still referenced</strong>: Foreign key constraints will block the drop, but application code will crash. Remove all code references first, then drop.</p></li><li><p><strong>Running large data migrations in the main transaction</strong>: Updating millions of rows in a single transaction locks the table and can cause timeouts. Batch your updates (1000-5000 rows at a time) with small pauses between batches.</p></li></ul></blockquote><p>A useful rule of thumb: if a migration cannot be reversed with a simple &#8220;undo&#8221; migration, it is dangerous and needs the expand-contract treatment.</p><h5><strong>Zero-downtime deployments in Kubernetes</strong></h5><p>Now that we know how to handle database changes safely, let&#8217;s look at the other half of the puzzle: deploying application code without dropping requests. Kubernetes gives you several mechanisms for this.</p><p><strong>Rolling updates</strong></p><p>The default Kubernetes deployment strategy is <code>RollingUpdate</code>. It gradually replaces old pods with new pods, ensuring that some pods are always available to serve traffic.</p><pre><code>apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1        # At most 1 extra pod during rollout
      maxUnavailable: 0   # Never have fewer than 3 healthy pods
  selector:
    matchLabels:
      app: myapp
  template:
    metadata:
      labels:
        app: myapp
    spec:
      containers:
        - name: myapp
          image: myapp:v2
          ports:
            - containerPort: 3000</code></pre><p>With <code>maxSurge: 1</code> and <code>maxUnavailable: 0</code>, Kubernetes will:</p><blockquote><ul><li><p><strong>Step 1</strong>: Create 1 new pod (v2). Now you have 3 old + 1 new = 4 pods total.</p></li><li><p><strong>Step 2</strong>: Wait until the new pod passes its readiness probe.</p></li><li><p><strong>Step 3</strong>: Terminate 1 old pod (v1). Now you have 2 old + 1 new = 3 pods.</p></li><li><p><strong>Step 4</strong>: Create another new pod (v2). Now you have 2 old + 2 new = 4 pods.</p></li><li><p><strong>Repeat</strong> until all pods are running v2.</p></li></ul></blockquote><p>During this process, both v1 and v2 pods are serving traffic. This is exactly why your database schema must be compatible with both versions.</p><p><strong>The Recreate strategy</strong></p><p>The <code>Recreate</code> strategy kills all old pods before creating new ones. This means downtime, so you should only use it when your application cannot run two versions simultaneously (for example, if it holds an exclusive lock on a resource).</p><pre><code>strategy:
  type: Recreate</code></pre><p>For almost all web applications, <code>RollingUpdate</code> is what you want.</p><h5><strong>Health checks: liveness, readiness, and startup probes</strong></h5><p>Probes are how Kubernetes knows if your pod is healthy. There are three types, and each one serves a different purpose:</p><blockquote><ul><li><p><strong>Readiness probe</strong>: &#8220;Is this pod ready to receive traffic?&#8221; Kubernetes only sends traffic to pods that pass their readiness probe. During a deployment, new pods will not receive traffic until they are ready. This is the most important probe for zero-downtime deployments.</p></li><li><p><strong>Liveness probe</strong>: &#8220;Is this pod still alive?&#8221; If a pod fails its liveness probe, Kubernetes restarts it. Use this to recover from deadlocks or stuck processes. Be careful: if your liveness probe is too aggressive, Kubernetes will restart pods that are just slow, creating a crash loop.</p></li><li><p><strong>Startup probe</strong>: &#8220;Has this pod finished starting up?&#8221; This is for applications with a slow startup (loading large caches, running migrations). The startup probe runs first, and liveness/readiness probes do not start until it passes.</p></li></ul></blockquote><p>Here is how to configure all three:</p><pre><code>containers:
  - name: myapp
    image: myapp:v2
    ports:
      - containerPort: 3000
    readinessProbe:
      httpGet:
        path: /health/ready
        port: 3000
      initialDelaySeconds: 5
      periodSeconds: 10
      failureThreshold: 3
    livenessProbe:
      httpGet:
        path: /health/live
        port: 3000
      initialDelaySeconds: 15
      periodSeconds: 20
      failureThreshold: 3
    startupProbe:
      httpGet:
        path: /health/started
        port: 3000
      failureThreshold: 30
      periodSeconds: 10</code></pre><p>And here is what the health check endpoints look like in your TypeScript API:</p><pre><code>// Health check endpoints
app.get("/health/live", (req, res) =&gt; {
  // Liveness: is the process running?
  // Keep this simple. If this endpoint responds, the process is alive.
  res.status(200).json({ status: "alive" });
});

app.get("/health/ready", async (req, res) =&gt; {
  // Readiness: can this pod serve traffic?
  // Check that all dependencies are reachable.
  try {
    await prisma.$queryRaw`SELECT 1`;  // database is reachable
    res.status(200).json({ status: "ready" });
  } catch (error) {
    res.status(503).json({ status: "not ready", error: "database unreachable" });
  }
});

app.get("/health/started", (req, res) =&gt; {
  // Startup: has the app finished initializing?
  if (appIsInitialized) {
    res.status(200).json({ status: "started" });
  } else {
    res.status(503).json({ status: "starting" });
  }
});</code></pre><p>A common mistake is making the liveness probe too strict. If your liveness probe checks the database, and the database has a brief network hiccup, Kubernetes will restart all your pods at once, making the situation much worse. Keep liveness probes simple (just &#8220;is the process running?&#8221;) and use readiness probes for dependency checks.</p><h5><strong>Graceful shutdown: preStop hooks and connection draining</strong></h5><p>When Kubernetes terminates a pod during a rolling update, it sends a <code>SIGTERM</code> signal. Your application should catch this signal and stop accepting new requests while finishing in-flight requests. But there is a race condition: Kubernetes removes the pod from the service endpoints at the same time it sends <code>SIGTERM</code>, and the endpoint removal takes a moment to propagate. During that window, traffic can still be routed to a pod that is shutting down.</p><p>The fix is a <code>preStop</code> hook that adds a small delay:</p><pre><code>containers:
  - name: myapp
    image: myapp:v2
    lifecycle:
      preStop:
        exec:
          command: ["sh", "-c", "sleep 10"]</code></pre><p>This tells Kubernetes to wait 10 seconds before sending <code>SIGTERM</code>. During those 10 seconds, the pod is removed from the service endpoints, so no new traffic is routed to it. After the sleep, <code>SIGTERM</code> is sent and the application can shut down gracefully.</p><p>In your TypeScript application, handle the shutdown signal:</p><pre><code>// Graceful shutdown handler
process.on("SIGTERM", () =&gt; {
  console.log("SIGTERM received. Starting graceful shutdown...");

  // Stop accepting new connections
  server.close(() =&gt; {
    console.log("HTTP server closed. Cleaning up...");

    // Close database connections
    prisma.$disconnect().then(() =&gt; {
      console.log("Database disconnected. Exiting.");
      process.exit(0);
    });
  });

  // Force exit after 30 seconds if graceful shutdown hangs
  setTimeout(() =&gt; {
    console.error("Graceful shutdown timed out. Forcing exit.");
    process.exit(1);
  }, 30000);
});</code></pre><p>Also, set <code>terminationGracePeriodSeconds</code> on the pod spec to give your application enough time to drain. The default is 30 seconds, but adjust it based on how long your longest requests take:</p><pre><code>spec:
  terminationGracePeriodSeconds: 60
  containers:
    - name: myapp
      # ...</code></pre><h5><strong>PodDisruptionBudgets</strong></h5><p>A PodDisruptionBudget (PDB) tells Kubernetes how many pods must remain available during voluntary disruptions like node drains, cluster upgrades, or autoscaler scale-downs. Without a PDB, Kubernetes could drain all your nodes at once during a cluster upgrade, taking down every pod simultaneously.</p><pre><code>apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: myapp-pdb
spec:
  minAvailable: 2    # At least 2 pods must always be running
  selector:
    matchLabels:
      app: myapp</code></pre><p>You can also use <code>maxUnavailable</code> instead of <code>minAvailable</code>:</p><pre><code>spec:
  maxUnavailable: 1   # At most 1 pod can be down at a time</code></pre><p>For a deployment with 3 replicas, <code>minAvailable: 2</code> and <code>maxUnavailable: 1</code> are equivalent. Use whichever reads more clearly for your team.</p><h5><strong>Blue-green deployments</strong></h5><p>Blue-green deployments take a different approach from rolling updates. Instead of gradually replacing pods, you run two complete environments simultaneously: the &#8220;blue&#8221; environment (current production) and the &#8220;green&#8221; environment (new version). Once the green environment is validated, you switch traffic from blue to green in one step.</p><p>Here is how to implement blue-green with Kubernetes services:</p><pre><code># Blue deployment (current production, running v1)
apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp-blue
spec:
  replicas: 3
  selector:
    matchLabels:
      app: myapp
      version: blue
  template:
    metadata:
      labels:
        app: myapp
        version: blue
    spec:
      containers:
        - name: myapp
          image: myapp:v1
---
# Green deployment (new version, running v2)
apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp-green
spec:
  replicas: 3
  selector:
    matchLabels:
      app: myapp
      version: green
  template:
    metadata:
      labels:
        app: myapp
        version: green
    spec:
      containers:
        - name: myapp
          image: myapp:v2
---
# Service (points to blue initially)
apiVersion: v1
kind: Service
metadata:
  name: myapp
spec:
  selector:
    app: myapp
    version: blue    # Change to "green" to switch traffic
  ports:
    - port: 80
      targetPort: 3000</code></pre><p>To switch traffic, you update the service selector from <code>version: blue</code> to <code>version: green</code>. All traffic moves at once. If something is wrong, you switch back to blue.</p><blockquote><ul><li><p><strong>When to use blue-green</strong>: When you need instant rollback (one label change), when you want to validate the new version with real traffic patterns before committing, or when your application cannot tolerate mixed versions.</p></li><li><p><strong>When to stick with rolling updates</strong>: For most standard deployments. Blue-green requires double the resources (both environments run simultaneously) and adds operational complexity.</p></li></ul></blockquote><h5><strong>Rollback strategies</strong></h5><p>No matter how careful you are, things will go wrong. Having a rollback plan is not optional. Here are the three main strategies:</p><p><strong>Kubernetes rollback</strong></p><p>Kubernetes keeps a history of your deployments. You can roll back to a previous version with a single command:</p><pre><code># See rollout history
kubectl rollout history deployment/myapp

# Roll back to the previous version
kubectl rollout undo deployment/myapp

# Roll back to a specific revision
kubectl rollout undo deployment/myapp --to-revision=3

# Watch the rollback progress
kubectl rollout status deployment/myapp</code></pre><p>This only rolls back the application code (the container image). It does not roll back database migrations. If your migration was additive (adding a column), the old code simply ignores the new column, and there is nothing to roll back. If your migration was destructive (dropping a column), you need a database rollback.</p><p><strong>Database rollback</strong></p><p>Prisma does not have a built-in &#8220;undo last migration&#8221; command for production. In production, you write a new migration that reverses the change:</p><pre><code># In development, you can reset (destroys all data)
npx prisma migrate reset

# In production, create a new "undo" migration
npx prisma migrate dev --name undo_add_phone_to_users</code></pre><p>The &#8220;undo&#8221; migration is just another migration that reverses the previous change:</p><pre><code>-- Undo migration: remove the phone column
ALTER TABLE "User" DROP COLUMN "phone";</code></pre><p>This is another reason to prefer additive migrations. Adding a column is easy to undo (drop it). Dropping a column is impossible to undo (the data is gone). If you follow the expand-contract pattern, your &#8220;undo&#8221; is always just &#8220;drop the column you added.&#8221;</p><p><strong>Feature flags as an alternative to rollbacks</strong></p><p>Instead of rolling back code or database changes, you can use feature flags to disable the new functionality without changing the deployed code:</p><pre><code>// Feature flag check
app.get("/api/orders", async (req, res) =&gt; {
  const orders = await prisma.order.findMany({
    include: {
      user: true,
    },
  });

  if (featureFlags.isEnabled("show-order-currency")) {
    // New behavior: include currency field
    return res.json(orders.map(o =&gt; ({
      ...o,
      currency: o.currency ?? "USD",
    })));
  }

  // Old behavior: no currency field
  return res.json(orders);
});</code></pre><p>Feature flags let you decouple deployment from release. You deploy the code (including the migration), but the new feature is behind a flag. If something goes wrong, you flip the flag off. No rollback needed, no redeployment, no database undo.</p><h5><strong>Practical example: adding a column under production traffic</strong></h5><p>Let&#8217;s put it all together with a real scenario. We need to add a <code>currency</code> column to the <code>Order</code> table. The API is serving traffic, and we cannot afford any downtime.</p><p><strong>Step 1: Write the migration</strong></p><p>Update the Prisma schema:</p><pre><code>model Order {
  id        Int      @id @default(autoincrement())
  amount    Float
  status    String   @default("pending")
  currency  String   @default("USD")   // new column with a default
  userId    Int
  user      User     @relation(fields: [userId], references: [id])
  createdAt DateTime @default(now())
}</code></pre><p>Generate and review the migration:</p><pre><code>npx prisma migrate dev --name add_currency_to_orders</code></pre><p>Generated SQL:</p><pre><code>ALTER TABLE "Order" ADD COLUMN "currency" TEXT NOT NULL DEFAULT 'USD';</code></pre><p>This is safe because the column has a default value, so existing rows get <code>'USD'</code> automatically, and old application code that does not know about the column will simply ignore it.</p><p><strong>Step 2: Update the application code</strong></p><p>Update the API to use the new column:</p><pre><code>// Updated order creation endpoint
app.post("/api/orders", async (req, res) =&gt; {
  const { amount, userId, currency } = req.body;

  const order = await prisma.order.create({
    data: {
      amount,
      userId,
      currency: currency ?? "USD",   // use provided currency or default
    },
  });

  res.status(201).json(order);
});</code></pre><p><strong>Step 3: Run the migration in CI/CD</strong></p><p>Add a migration step to your CI/CD pipeline that runs before the application deployment:</p><pre><code># In your GitHub Actions workflow
jobs:
  migrate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - name: Setup Node
        uses: actions/setup-node@v4
        with:
          node-version: "20"

      - name: Install dependencies
        run: npm ci

      - name: Run database migrations
        run: npx prisma migrate deploy
        env:
          DATABASE_URL: ${{ secrets.DATABASE_URL }}

  deploy:
    needs: migrate    # Deploy only after migrations succeed
    runs-on: ubuntu-latest
    steps:
      - name: Update deployment image
        run: |
          kubectl set image deployment/myapp \
            myapp=myapp:${{ github.sha }}

      - name: Wait for rollout
        run: kubectl rollout status deployment/myapp --timeout=300s</code></pre><p><strong>Step 4: Verify</strong></p><p>After the deploy, verify everything is working:</p><pre><code># Check the migration was applied
npx prisma migrate status

# Test the API
curl -X POST https://api.example.com/api/orders \
  -H "Content-Type: application/json" \
  -d '{"amount": 29.99, "userId": 1, "currency": "EUR"}'

# Verify the response includes the currency
curl https://api.example.com/api/orders/1</code></pre><p>Because we used an additive change with a default value, the entire process was zero-downtime. Old pods that do not know about the <code>currency</code> column kept serving traffic while new pods rolled out. No conflicts, no errors, no interruption.</p><h5><strong>Migration checklist</strong></h5><p>Before running any migration in production, go through this checklist:</p><blockquote><ul><li><p><strong>Is the migration additive?</strong> Adding columns (nullable or with defaults), adding tables, and adding indexes are safe. Everything else needs extra care.</p></li><li><p><strong>Can old code work with the new schema?</strong> During a rolling update, old and new code run simultaneously. Make sure the old code will not break.</p></li><li><p><strong>Can new code work with the old schema?</strong> If the migration fails or is delayed, can the new code still function?</p></li><li><p><strong>Have you tested the migration on a copy of production data?</strong> Your dev database has 100 rows. Production has 10 million. What takes 1 second in dev might take 10 minutes in production.</p></li><li><p><strong>Do you have a rollback plan?</strong> What SQL would you run to undo this migration? Write it down before you deploy.</p></li><li><p><strong>Are you using <code>CONCURRENTLY</code> for index creation?</strong> On large tables, index creation locks the table. Use <code>CREATE INDEX CONCURRENTLY</code> in PostgreSQL.</p></li><li><p><strong>Are you batching large data migrations?</strong> Do not update millions of rows in a single transaction. Batch them.</p></li></ul></blockquote><h5><strong>What comes next</strong></h5><p>We now know how to make database changes safely, deploy application code without downtime, and roll back when things go wrong. In the next article, we will explore security in the CI/CD pipeline: scanning for vulnerabilities, managing secrets, and hardening your deployment process.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: CI/CD, The Complete Pipeline]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-the-complete-pipeline</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-the-complete-pipeline</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vp5O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vp5O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vp5O!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!vp5O!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!vp5O!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!vp5O!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vp5O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043123?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vp5O!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!vp5O!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!vp5O!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!vp5O!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb2b9f372-2ab1-495a-8ae9-b58cd1c6852d_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article sixteen of the DevOps from Zero to Hero series. Over the past fifteen articles we have covered everything from writing a TypeScript API, to version control, testing, CI, infrastructure as code, Kubernetes, Helm, secrets, and more. Each piece solved a specific problem, but we have not yet stitched them all together into one cohesive, end-to-end pipeline.</p><p>That changes now. In this article we are going to build a complete CI/CD pipeline that takes your code from a pull request all the way to production. Not a toy example. A real, multi-job GitHub Actions workflow that lints, tests, builds, deploys to staging, runs smoke tests, waits for manual approval, and then promotes to production. We will also cover deployment strategies, rollback procedures, and best practices for keeping your pipeline fast and reliable.</p><p>If you have been following the series, think of this article as the glue that connects everything. If you are jumping in fresh, do not worry. We will explain each piece as we go.</p><p>Let&#8217;s get into it.</p><h5><strong>The pipeline philosophy</strong></h5><p>Before we write a single line of YAML, let&#8217;s establish the principles that drive a good CI/CD pipeline:</p><blockquote><ul><li><p><strong>Every commit to main should be deployable</strong>: If something is in main, it has been linted, tested, and built. It is ready to ship. If it is not ready, it should not be in main.</p></li><li><p><strong>Environments are gates, not destinations</strong>: Staging exists to validate, not to accumulate. Code should flow through staging quickly, not sit there for weeks. Production is the destination.</p></li><li><p><strong>Fail fast, fail loud</strong>: If something is broken, you want to know in seconds, not minutes. Put the cheapest checks first (lint, format) and the expensive ones later (integration tests, builds).</p></li><li><p><strong>Automation over manual processes</strong>: Every manual step is a step that can be forgotten, done wrong, or skipped under pressure. Automate everything except the final production approval.</p></li><li><p><strong>Reproducibility</strong>: Your pipeline should produce the same result whether you run it today or three months from now. Pin your versions, cache your dependencies, and use immutable artifacts.</p></li></ul></blockquote><p>These are not abstract ideals. They are engineering decisions that prevent outages, reduce toil, and let you ship with confidence. Every design choice in the pipeline we are about to build traces back to one of these principles.</p><h5><strong>Pipeline stages overview</strong></h5><p>Our pipeline will have seven stages, organized into three phases:</p><pre><code>Phase 1: Validate (on every PR and push to main)
  &#9500;&#9472;&#9472; Lint       -&gt; ESLint, Prettier, type checking
  &#9492;&#9472;&#9472; Test       -&gt; Unit tests, integration tests, coverage

Phase 2: Build and Deploy to Staging (on push to main only)
  &#9500;&#9472;&#9472; Build      -&gt; Docker image build and push to registry
  &#9500;&#9472;&#9472; Deploy     -&gt; Deploy to staging namespace via ArgoCD
  &#9492;&#9472;&#9472; Smoke Test -&gt; Health check and API tests against staging

Phase 3: Promote to Production (manual approval)
  &#9500;&#9472;&#9472; Approve    -&gt; Manual approval gate via GitHub Environments
  &#9492;&#9472;&#9472; Deploy     -&gt; Deploy to production namespace</code></pre><p>Phase 1 runs on every pull request and every push to main. It is your safety net. Phase 2 only runs on pushes to main (merged PRs) because you do not want to deploy feature branches to staging. Phase 3 requires a human to click &#8220;Approve&#8221; before code reaches production. This is the one manual step we keep on purpose, because deploying to production should be a conscious decision.</p><h5><strong>GitHub Actions environments</strong></h5><p>GitHub Actions has a feature called Environments that gives you exactly what we need: environment-specific secrets, protection rules, and deployment history. Let&#8217;s set them up.</p><p>Go to your repository on GitHub, then Settings, then Environments. Create two environments:</p><blockquote><ul><li><p><strong>staging</strong>: No protection rules needed. Deployments here should be automatic after the build passes.</p></li><li><p><strong>production</strong>: Add a &#8220;Required reviewers&#8221; protection rule. Pick one or more team members who must approve before a deployment can proceed.</p></li></ul></blockquote><p>You can also add a &#8220;Wait timer&#8221; to production if you want a mandatory cooldown period between staging and production deploys. Some teams set this to 15 minutes to give smoke tests extra time to surface issues.</p><h5><strong>Environment-specific secrets and variables</strong></h5><p>Each environment can have its own secrets and variables. This is how you handle the fact that staging and production use different clusters, namespaces, databases, and API keys without littering your workflow with <code>if</code> conditionals.</p><p>Here is what you would typically configure:</p><pre><code>Repository secrets (shared):
  REGISTRY_USERNAME    -&gt; your container registry username
  REGISTRY_PASSWORD    -&gt; your container registry token

Staging environment secrets:
  KUBE_CONFIG          -&gt; kubeconfig for your staging cluster
  DATABASE_URL         -&gt; staging database connection string
  ARGOCD_AUTH_TOKEN    -&gt; ArgoCD token for staging

Staging environment variables:
  KUBE_NAMESPACE       -&gt; staging
  APP_URL              -&gt; https://staging.myapp.example.com

Production environment secrets:
  KUBE_CONFIG          -&gt; kubeconfig for your production cluster
  DATABASE_URL         -&gt; production database connection string
  ARGOCD_AUTH_TOKEN    -&gt; ArgoCD token for production

Production environment variables:
  KUBE_NAMESPACE       -&gt; production
  APP_URL              -&gt; https://myapp.example.com</code></pre><p>When a job specifies <code>environment: staging</code>, it can only access the staging secrets and variables. When it specifies <code>environment: production</code>, it gets the production ones. This isolation prevents the worst kind of mistake: accidentally running a production migration against the staging database, or vice versa.</p><p>To configure these, go to Settings, then Environments, click on the environment, and add your secrets and variables there. They work exactly like repository-level secrets but are scoped to the environment.</p><h5><strong>The complete workflow</strong></h5><p>Here is the full pipeline. We will go through each job in detail after, but first, see the big picture:</p><pre><code>name: CI/CD Pipeline

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

env:
  REGISTRY: ghcr.io
  IMAGE_NAME: ${{ github.repository }}

permissions:
  contents: read
  packages: write

jobs:
  # Phase 1: Validate
  lint:
    name: Lint
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"

      - run: npm ci

      - name: Run ESLint
        run: npx eslint .

      - name: Check formatting
        run: npx prettier --check .

      - name: Type check
        run: npx tsc --noEmit

  test:
    name: Test
    runs-on: ubuntu-latest
    services:
      postgres:
        image: postgres:16
        env:
          POSTGRES_USER: test
          POSTGRES_PASSWORD: test
          POSTGRES_DB: myapp_test
        ports:
          - 5432:5432
        options: &gt;-
          --health-cmd pg_isready
          --health-interval 10s
          --health-timeout 5s
          --health-retries 5
    env:
      DATABASE_URL: postgres://test:test@localhost:5432/myapp_test
    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"

      - run: npm ci

      - name: Run tests with coverage
        run: npm test -- --coverage

      - name: Upload coverage
        if: github.event_name == 'push'
        uses: actions/upload-artifact@v4
        with:
          name: coverage-report
          path: coverage/
          retention-days: 14

  # Phase 2: Build and Deploy to Staging
  build:
    name: Build and Push Image
    runs-on: ubuntu-latest
    needs: [lint, test]
    if: github.event_name == 'push' &amp;&amp; github.ref == 'refs/heads/main'
    outputs:
      image-tag: ${{ steps.meta.outputs.tags }}
      image-digest: ${{ steps.build.outputs.digest }}
    steps:
      - uses: actions/checkout@v4

      - uses: docker/setup-buildx-action@v3

      - name: Log in to registry
        uses: docker/login-action@v3
        with:
          registry: ${{ env.REGISTRY }}
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - name: Extract metadata
        id: meta
        uses: docker/metadata-action@v5
        with:
          images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}
          tags: |
            type=sha,prefix=
            type=raw,value=latest

      - name: Build and push
        id: build
        uses: docker/build-push-action@v6
        with:
          context: .
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          cache-from: type=gha
          cache-to: type=gha,mode=max

  deploy-staging:
    name: Deploy to Staging
    runs-on: ubuntu-latest
    needs: [build]
    environment: staging
    steps:
      - uses: actions/checkout@v4

      - name: Install ArgoCD CLI
        run: |
          curl -sSL -o argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
          chmod +x argocd
          sudo mv argocd /usr/local/bin/

      - name: Deploy to staging
        env:
          ARGOCD_SERVER: ${{ vars.ARGOCD_SERVER }}
          ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
        run: |
          argocd app set myapp-staging \
            --parameter image.tag=${{ github.sha }} \
            --grpc-web

          argocd app sync myapp-staging \
            --grpc-web \
            --timeout 300

          argocd app wait myapp-staging \
            --grpc-web \
            --timeout 300 \
            --health

  smoke-test:
    name: Smoke Tests
    runs-on: ubuntu-latest
    needs: [deploy-staging]
    environment: staging
    steps:
      - uses: actions/checkout@v4

      - name: Wait for deployment to stabilize
        run: sleep 30

      - name: Health check
        run: |
          for i in $(seq 1 10); do
            STATUS=$(curl -s -o /dev/null -w "%{http_code}" \
              "${{ vars.APP_URL }}/health")
            if [ "$STATUS" = "200" ]; then
              echo "Health check passed on attempt $i"
              exit 0
            fi
            echo "Attempt $i: got $STATUS, retrying in 10s..."
            sleep 10
          done
          echo "Health check failed after 10 attempts"
          exit 1

      - name: API smoke test
        run: |
          RESPONSE=$(curl -s -w "\n%{http_code}" \
            "${{ vars.APP_URL }}/api/v1/status")
          BODY=$(echo "$RESPONSE" | head -n -1)
          STATUS=$(echo "$RESPONSE" | tail -n 1)

          echo "Status: $STATUS"
          echo "Body: $BODY"

          if [ "$STATUS" != "200" ]; then
            echo "API smoke test failed with status $STATUS"
            exit 1
          fi

          echo "API smoke test passed"

      - name: Run E2E tests against staging
        env:
          BASE_URL: ${{ vars.APP_URL }}
        run: |
          npm ci
          npx playwright test tests/e2e/smoke.spec.ts

  # Phase 3: Promote to Production
  deploy-production:
    name: Deploy to Production
    runs-on: ubuntu-latest
    needs: [smoke-test]
    environment: production
    steps:
      - uses: actions/checkout@v4

      - name: Install ArgoCD CLI
        run: |
          curl -sSL -o argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
          chmod +x argocd
          sudo mv argocd /usr/local/bin/

      - name: Deploy to production
        env:
          ARGOCD_SERVER: ${{ vars.ARGOCD_SERVER }}
          ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
        run: |
          argocd app set myapp-production \
            --parameter image.tag=${{ github.sha }} \
            --grpc-web

          argocd app sync myapp-production \
            --grpc-web \
            --timeout 300

          argocd app wait myapp-production \
            --grpc-web \
            --timeout 300 \
            --health

      - name: Verify production deployment
        run: |
          for i in $(seq 1 10); do
            STATUS=$(curl -s -o /dev/null -w "%{http_code}" \
              "${{ vars.APP_URL }}/health")
            if [ "$STATUS" = "200" ]; then
              echo "Production health check passed"
              exit 0
            fi
            echo "Attempt $i: got $STATUS, retrying in 10s..."
            sleep 10
          done
          echo "Production health check failed"
          exit 1</code></pre><p>That is a lot of YAML, so let&#8217;s break it down piece by piece.</p><h5><strong>Phase 1: Validate</strong></h5><p>The lint and test jobs run in parallel on every push and pull request. They are the cheapest and fastest checks, so they go first.</p><p>The lint job runs three checks: ESLint for code quality, Prettier for formatting, and the TypeScript compiler for type safety. If any of these fail, the pipeline stops. There is no point building a Docker image for code that does not compile.</p><p>The test job spins up a PostgreSQL service container. GitHub Actions lets you define services alongside your job, and they are available on <code>localhost</code> just like a local database. The tests run with coverage enabled, and the coverage report is uploaded as an artifact for later review.</p><p>Notice that lint and test have no dependency on each other. They run in parallel by default, which means the validate phase takes as long as the slower of the two, not the sum of both.</p><h5><strong>Phase 2: Build and deploy to staging</strong></h5><p>The build job only runs on pushes to main (not on pull requests) and only after both lint and test pass. This is controlled by the <code>needs: [lint, test]</code> dependency and the <code>if</code> conditional.</p><p>We use Docker Buildx with GitHub Actions cache (<code>cache-from: type=gha</code>). This means subsequent builds reuse cached layers, which can cut build time from minutes to seconds. The image is tagged with the Git SHA and pushed to GitHub Container Registry (GHCR).</p><p>The deploy-staging job uses the ArgoCD CLI to update the image tag and sync the application. ArgoCD then handles the actual Kubernetes deployment: it updates the deployment manifest, waits for the new pods to be healthy, and reports back. The <code>argocd app wait</code> command blocks until the deployment is fully rolled out and healthy, so the pipeline knows whether the deploy succeeded or failed.</p><p>If you are not using ArgoCD, you can replace this with <code>kubectl</code> commands:</p><pre><code>      - name: Deploy to staging with kubectl
        run: |
          echo "${{ secrets.KUBE_CONFIG }}" | base64 -d &gt; kubeconfig
          export KUBECONFIG=kubeconfig

          kubectl set image deployment/myapp \
            myapp=${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}:${{ github.sha }} \
            -n staging

          kubectl rollout status deployment/myapp \
            -n staging \
            --timeout=300s

          rm kubeconfig</code></pre><p>The key point is the same: update the image, then wait for the rollout to finish before moving on.</p><h5><strong>Smoke tests in detail</strong></h5><p>The smoke test job is the gatekeeper between staging and production. It answers one question: is the thing we just deployed actually working?</p><p>We run three levels of smoke tests:</p><blockquote><ul><li><p><strong>Health check</strong>: A simple HTTP request to <code>/health</code>. If the server is not responding, everything else is irrelevant. We retry up to 10 times with 10-second intervals because deployments can take a moment to stabilize.</p></li><li><p><strong>API smoke test</strong>: A request to a real API endpoint. This validates that the application is not just running but actually serving requests correctly. We check both the status code and that the response body is valid.</p></li><li><p><strong>E2E smoke test</strong>: A Playwright test that loads the application in a browser and performs a few critical user flows. This catches issues that API-level tests miss, like broken JavaScript bundles or misconfigured CDN paths.</p></li></ul></blockquote><p>You do not need all three levels on day one. Start with just the health check. Add the API test when you have an API. Add the E2E test when you have Playwright set up. The important thing is to have something that validates the deployment before you promote to production.</p><p>Here is a minimal Playwright smoke test:</p><pre><code>import { test, expect } from "@playwright/test";

const BASE_URL = process.env.BASE_URL || "http://localhost:3000";

test.describe("Smoke Tests", () =&gt; {
  test("homepage loads successfully", async ({ page }) =&gt; {
    const response = await page.goto(BASE_URL);
    expect(response?.status()).toBe(200);
    await expect(page.locator("h1")).toBeVisible();
  });

  test("API returns valid response", async ({ request }) =&gt; {
    const response = await request.get(`${BASE_URL}/api/v1/status`);
    expect(response.status()).toBe(200);

    const body = await response.json();
    expect(body).toHaveProperty("status", "ok");
  });

  test("login page renders", async ({ page }) =&gt; {
    await page.goto(`${BASE_URL}/login`);
    await expect(page.locator('input[name="email"]')).toBeVisible();
    await expect(page.locator('button[type="submit"]')).toBeVisible();
  });
});</code></pre><p>Keep smoke tests fast. They should run in under a minute. If you need comprehensive E2E coverage, run that in a separate workflow. Smoke tests are about confidence, not completeness.</p><h5><strong>Production promotion and manual approval</strong></h5><p>The deploy-production job has <code>environment: production</code>, which triggers the protection rules you configured earlier. When the pipeline reaches this job, it pauses and shows a &#8220;Review deployments&#8221; button in the GitHub Actions UI. The required reviewers you configured get a notification, and the pipeline waits until one of them clicks &#8220;Approve.&#8221;</p><p>This is intentional. Production deployments should be a deliberate decision. The approval step gives your team a moment to ask: did the smoke tests look good? Are there any known issues? Is this a good time to deploy (not Friday afternoon)?</p><p>Once approved, the production deploy follows the same pattern as staging: update the image tag, sync with ArgoCD, wait for the rollout, and verify with a health check.</p><p>You might be wondering why we do not run the full smoke test suite against production. Some teams do, and that is fine. But there is a tradeoff: running tests against production means your tests can fail due to production-specific issues (rate limiting, real data edge cases), and a test failure after deploy can cause confusion about whether the deploy itself failed. A simple health check is usually enough for the production verification step.</p><h5><strong>Deployment strategies</strong></h5><p>The pipeline we built uses the default Kubernetes deployment strategy: rolling update. But it is worth understanding the alternatives and when to use them.</p><p><strong>Rolling update (default)</strong></p><p>This is what Kubernetes does out of the box. It gradually replaces old pods with new pods, one at a time (or in batches). At any point during the rollout, some pods are running the old version and some are running the new version.</p><pre><code>apiVersion: apps/v1
kind: Deployment
metadata:
  name: myapp
spec:
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxSurge: 1
      maxUnavailable: 0
  replicas: 3
  template:
    spec:
      containers:
        - name: myapp
          image: ghcr.io/myorg/myapp:abc123
          readinessProbe:
            httpGet:
              path: /health
              port: 3000
            initialDelaySeconds: 5
            periodSeconds: 10</code></pre><blockquote><ul><li><p><strong>maxSurge: 1</strong> means Kubernetes can create one extra pod above the desired replica count during the rollout.</p></li><li><p><strong>maxUnavailable: 0</strong> means no pod is removed until its replacement is ready. This ensures zero downtime.</p></li><li><p><strong>readinessProbe</strong> tells Kubernetes when a new pod is ready to receive traffic. Without this, Kubernetes might send requests to a pod that is still starting up.</p></li></ul></blockquote><p>Rolling updates are the right choice for most applications. They are simple, zero-downtime, and well-supported by every Kubernetes distribution.</p><p><strong>Blue-green deployment</strong></p><p>In a blue-green deployment, you run two identical environments: blue (current production) and green (the new version). Traffic goes to blue while green is being deployed and tested. Once green is verified, you switch traffic from blue to green in one shot.</p><p>The advantage is that the switch is instantaneous and you can roll back by switching back to blue. The disadvantage is that you need double the resources during the deployment. In Kubernetes, you can implement blue-green by maintaining two deployments and switching the service selector:</p><pre><code># Deploy the new version as "green"
kubectl set image deployment/myapp-green \
  myapp=ghcr.io/myorg/myapp:new-version -n production

kubectl rollout status deployment/myapp-green -n production

# Switch traffic from blue to green
kubectl patch service myapp \
  -p '{"spec":{"selector":{"version":"green"}}}' -n production</code></pre><p><strong>Canary deployment</strong></p><p>A canary deployment routes a small percentage of traffic (say 5%) to the new version while the majority continues hitting the old version. You monitor error rates and latency for the canary, and if everything looks good, you gradually increase the traffic split until 100% goes to the new version.</p><p>Canary deployments are powerful but require a service mesh (like Istio or Linkerd) or an ingress controller that supports traffic splitting. They are more complex to set up but give you the safest possible production rollout for high-traffic applications.</p><p>For this series, we will stick with rolling updates. They cover the vast majority of use cases, and you can always adopt blue-green or canary later when your needs grow.</p><h5><strong>Rollback strategies</strong></h5><p>Things go wrong. A deploy passes all tests but a subtle bug appears under real traffic. You need to get back to a known-good state fast. Here are your options:</p><p><strong>Option 1: Git revert and push</strong></p><p>This is the simplest and most reliable approach. You revert the commit that caused the problem, push to main, and the pipeline redeploys the previous version automatically.</p><pre><code># Find the commit that caused the issue
git log --oneline -5

# Revert it
git revert HEAD

# Push to main, which triggers the pipeline
git push origin main</code></pre><p>The advantage of this approach is that it goes through the full pipeline: lint, test, build, staging, smoke test, production. You know the reverted version works because it was validated at every stage. The downside is that it takes as long as a normal deployment (5-15 minutes depending on your pipeline).</p><p><strong>Option 2: ArgoCD rollback</strong></p><p>If you are using ArgoCD, you can roll back to a previous sync directly:</p><pre><code># List the sync history
argocd app history myapp-production

# Roll back to a specific revision
argocd app rollback myapp-production &lt;revision-number&gt;</code></pre><p>This is faster than a git revert because it skips the build step. ArgoCD simply redeploys the previous manifests. However, it creates a drift between your Git state and what is running in the cluster. You should still create a git revert afterwards to keep Git as the source of truth.</p><p><strong>Option 3: kubectl rollout undo</strong></p><p>Kubernetes keeps a history of deployments, and you can roll back with a single command:</p><pre><code># Roll back to the previous version
kubectl rollout undo deployment/myapp -n production

# Or roll back to a specific revision
kubectl rollout history deployment/myapp -n production
kubectl rollout undo deployment/myapp -n production --to-revision=3</code></pre><p>Like the ArgoCD rollback, this is fast but creates drift from Git. Use it for emergencies, then follow up with a proper git revert.</p><p>The recommendation is: for planned rollbacks, use git revert. For emergencies, use kubectl rollout undo or ArgoCD rollback, then git revert as a follow-up. Either way, Git should always reflect what is actually running in production.</p><h5><strong>Pipeline best practices</strong></h5><p>Now that you have a working pipeline, here are the practices that keep it fast, reliable, and maintainable over time:</p><p><strong>Fail early</strong></p><p>Order your jobs from fastest to slowest. Lint takes seconds, tests take a minute, Docker builds take several minutes. If the code does not pass lint, there is no point waiting for a Docker build to finish. The <code>needs</code> keyword enforces this ordering.</p><p><strong>Parallelize where possible</strong></p><p>Lint and test do not depend on each other. Run them in parallel. If you have multiple test suites (unit, integration, E2E), split them into separate jobs that run simultaneously. Every minute you shave off the pipeline is a minute your team gets back on every single commit.</p><p><strong>Cache aggressively</strong></p><p>Cache everything that does not change between builds:</p><blockquote><ul><li><p><strong>npm dependencies</strong>: Use <code>actions/setup-node</code> with <code>cache: "npm"</code>. This caches the npm global store and restores it based on <code>package-lock.json</code>.</p></li><li><p><strong>Docker layers</strong>: Use BuildKit with <code>cache-from: type=gha</code> and <code>cache-to: type=gha,mode=max</code>. This stores and restores layer caches using GitHub&#8217;s cache backend.</p></li><li><p><strong>Test fixtures</strong>: If your tests download large fixtures, cache them with <code>actions/cache</code>.</p></li></ul></blockquote><p>Without caching, a typical pipeline takes 8-12 minutes. With caching, it can drop to 3-5 minutes.</p><p><strong>Keep it DRY</strong></p><p>If you have multiple repositories with similar pipelines, extract common steps into reusable workflows or composite actions:</p><pre><code># .github/workflows/reusable-deploy.yml
name: Deploy
on:
  workflow_call:
    inputs:
      environment:
        required: true
        type: string
      argocd-app:
        required: true
        type: string
    secrets:
      ARGOCD_AUTH_TOKEN:
        required: true

jobs:
  deploy:
    runs-on: ubuntu-latest
    environment: ${{ inputs.environment }}
    steps:
      - name: Install ArgoCD CLI
        run: |
          curl -sSL -o argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
          chmod +x argocd
          sudo mv argocd /usr/local/bin/

      - name: Deploy
        env:
          ARGOCD_SERVER: ${{ vars.ARGOCD_SERVER }}
          ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}
        run: |
          argocd app set ${{ inputs.argocd-app }} \
            --parameter image.tag=${{ github.sha }} \
            --grpc-web
          argocd app sync ${{ inputs.argocd-app }} \
            --grpc-web --timeout 300
          argocd app wait ${{ inputs.argocd-app }} \
            --grpc-web --timeout 300 --health</code></pre><p>Then call it from your main pipeline:</p><pre><code>  deploy-staging:
    needs: [build]
    uses: ./.github/workflows/reusable-deploy.yml
    with:
      environment: staging
      argocd-app: myapp-staging
    secrets:
      ARGOCD_AUTH_TOKEN: ${{ secrets.ARGOCD_AUTH_TOKEN }}</code></pre><p>This avoids duplicating deployment logic across staging and production jobs. When you need to change how deployments work, you change it in one place.</p><p><strong>Pin your action versions</strong></p><p>Always use specific versions (or commit SHAs) for actions, not <code>@main</code> or <code>@latest</code>. Third-party actions can change without warning, and a broken action version can break your pipeline across all repositories at once:</p><pre><code># Good: pinned to a specific version
- uses: actions/checkout@v4
- uses: docker/build-push-action@v6

# Bad: unpinned, can break without warning
- uses: actions/checkout@main
- uses: some-org/some-action@latest</code></pre><h5><strong>Monitoring your pipeline</strong></h5><p>A pipeline is only useful if you know how it is performing. GitHub Actions gives you several ways to monitor pipeline health:</p><blockquote><ul><li><p><strong>Workflow run history</strong>: Go to the Actions tab in your repository. You can see every run, filter by workflow, branch, or status, and drill into individual jobs and steps.</p></li><li><p><strong>Build time trends</strong>: Track how long your pipeline takes over time. If builds are getting slower, it usually means your test suite is growing without corresponding optimization, or your Docker cache is not working correctly.</p></li><li><p><strong>Failure rate</strong>: If your pipeline fails more than 10% of the time on legitimate code changes, something is flaky. Common culprits are network-dependent tests, race conditions, and service container startup timing.</p></li><li><p><strong>Status badges</strong>: Add a workflow status badge to your README so the team can see pipeline health at a glance.</p></li></ul></blockquote><p>You can add a status badge to your README with this markdown:</p><pre><code>![CI/CD](https://github.com/myorg/myapp/actions/workflows/ci-cd.yml/badge.svg)</code></pre><p>For more advanced monitoring, consider integrating with tools like Datadog CI Visibility or Grafana with the GitHub Actions exporter. These give you dashboards with build time percentiles, failure breakdowns by job, and alerts when build times exceed a threshold.</p><h5><strong>Putting it all together</strong></h5><p>Let&#8217;s recap what happens when a developer pushes a change through this pipeline:</p><blockquote><ul><li><p><strong>Developer opens a PR</strong>: Lint and test run automatically. The PR gets a green checkmark or a red X. Code review happens in parallel.</p></li><li><p><strong>PR is merged to main</strong>: Lint and test run again on the merged code. Then the build job creates a Docker image tagged with the commit SHA and pushes it to GHCR.</p></li><li><p><strong>Staging deploy</strong>: ArgoCD updates the staging deployment with the new image tag. The pipeline waits until the rollout is healthy.</p></li><li><p><strong>Smoke tests</strong>: Health check, API test, and E2E test run against staging. If any fail, the pipeline stops and the team is notified.</p></li><li><p><strong>Manual approval</strong>: A reviewer checks the staging deployment, confirms it looks good, and clicks &#8220;Approve&#8221; in the GitHub Actions UI.</p></li><li><p><strong>Production deploy</strong>: ArgoCD updates the production deployment. A final health check confirms the deployment is live.</p></li></ul></blockquote><p>The entire process, from merge to production, takes about 10-15 minutes. Most of that time is in the build and test stages. The actual deployment steps take less than a minute each.</p><p>If anything goes wrong, the pipeline stops at the failed step. No code reaches production unless it has passed every gate. And if something slips through, you can roll back with a git revert in under a minute.</p><h5><strong>What comes next</strong></h5><p>We now have a complete, end-to-end CI/CD pipeline that takes code from a pull request to production with automated validation at every stage. In the next article, we will look at monitoring and observability: how to know what your application is doing once it is running in production.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Observability in Kubernetes]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-observability</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-observability</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!elaJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!elaJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!elaJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!elaJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!elaJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!elaJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!elaJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043127?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!elaJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!elaJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!elaJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!elaJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbdb79b79-1b03-400f-a319-ee253de3bdb3_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article fifteen of the DevOps from Zero to Hero series. In the previous articles we deployed our TypeScript API to Kubernetes and packaged it with Helm. Everything is running, the pods are green, and life is good. But then someone asks: &#8220;Is the API actually healthy? How do we know if response times are getting worse? What happened at 3am when users started complaining?&#8221;</p><p>Without observability, you are flying blind. You deployed your app, but you have no idea what is happening inside it. Observability gives you the ability to understand the internal state of your system by examining the data it produces. It is the difference between &#8220;something is broken&#8221; and &#8220;the /orders endpoint is returning 500 errors because the database connection pool is exhausted.&#8221;</p><p>In this article we will cover the three pillars of observability (logs, metrics, and traces), set up Prometheus and Grafana on EKS using Helm, build a basic dashboard, instrument our TypeScript API with structured logging and a metrics endpoint, configure a simple alert, and walk through the observability workflow you will use during real incidents. This is a beginner-friendly introduction. If you want to go deeper into topics like SLO-based alerting, Loki for log aggregation, or advanced OpenTelemetry patterns, check out the <a href="https://segfault.pw/blog/sre-observability-deep-dive-traces-logs-and-metrics">SRE Observability Deep Dive</a> from the SRE series.</p><p>Let&#8217;s get into it.</p><h5><strong>The three pillars of observability</strong></h5><p>Observability is built on three types of telemetry data. Each one answers a different question, and you need all three to debug production issues effectively.</p><blockquote><ul><li><p><strong>Logs</strong>: Discrete events that tell you what happened. &#8220;Request abc123 failed with a 500 error at 14:32:05.&#8221; Logs give you the richest context because they can include arbitrary details like request bodies, stack traces, and user IDs.</p></li><li><p><strong>Metrics</strong>: Numerical measurements over time. &#8220;The API handled 150 requests per second with a p99 latency of 200ms.&#8221; Metrics are cheap to store, fast to query, and perfect for dashboards and alerts.</p></li><li><p><strong>Traces</strong>: The path a request takes through your system. &#8220;This request hit the API gateway, then the orders service, then the database, and the slow part was the database query.&#8221; Traces are essential when you have multiple services talking to each other.</p></li></ul></blockquote><p>Think of it this way: metrics tell you something is wrong, traces tell you where in the system it is wrong, and logs tell you why it is wrong. Here is the flow:</p><pre><code># The observability workflow during an incident:
#
# 1. ALERT (from metrics): "Error rate &gt; 5% for the last 5 minutes"
#    -&gt; You know SOMETHING is wrong
#
# 2. DASHBOARD (metrics): Check Grafana, see /orders endpoint has high error rate
#    -&gt; You know WHAT is wrong
#
# 3. TRACES: Find failing requests, see they all fail at the database call
#    -&gt; You know WHERE it is wrong
#
# 4. LOGS: Check the database service logs: "ERROR: too many connections"
#    -&gt; You know WHY it is wrong</code></pre><p>We will cover each pillar in detail, starting with logs because they are the most familiar.</p><h5><strong>Logs: structured logging</strong></h5><p>If you have ever used <code>console.log("something broke")</code> in production, you know the problem. When you have thousands of log lines flowing through your system, finding the relevant one is like searching for a needle in a haystack. Unstructured logs (plain text strings) are hard to search, hard to filter, and hard to aggregate.</p><p>Structured logging solves this by writing logs as JSON objects with consistent fields. Instead of:</p><pre><code>[2026-06-02 14:32:05] ERROR: Failed to process order 12345 for user john@example.com</code></pre><p>You write:</p><pre><code>{
  "timestamp": "2026-06-02T14:32:05.123Z",
  "level": "error",
  "message": "Failed to process order",
  "orderId": "12345",
  "userId": "john@example.com",
  "service": "orders-api",
  "traceId": "abc123def456",
  "duration_ms": 1523
}</code></pre><p>Now you can search for all errors related to a specific user, a specific order, or a specific trace. You can count how many errors happened per service. You can correlate logs with traces using the traceId field. This is the power of structured logging.</p><p><strong>Log levels</strong> define the severity of a log entry. Use them consistently:</p><blockquote><ul><li><p><strong>error</strong>: Something failed and needs attention. A request returned a 500, a database query timed out, an external API is unreachable.</p></li><li><p><strong>warn</strong>: Something unexpected happened but the system handled it. A retry succeeded, a cache miss occurred, a deprecated endpoint was called.</p></li><li><p><strong>info</strong>: Normal operations worth recording. A request was processed successfully, a user logged in, a background job completed.</p></li><li><p><strong>debug</strong>: Detailed information useful during development. Request payloads, SQL queries, internal state. Disable this in production unless you are actively debugging.</p></li></ul></blockquote><p>Let&#8217;s add structured logging to our TypeScript API using <code>pino</code>, which is the fastest JSON logger for Node.js:</p><pre><code># Install pino and the pretty-printer for local development
npm install pino pino-http
npm install -D pino-pretty</code></pre><pre><code>// src/logger.ts
import pino from "pino";

const logger = pino({
  level: process.env.LOG_LEVEL || "info",
  // In production, output raw JSON. Locally, use pino-pretty for readability.
  transport:
    process.env.NODE_ENV !== "production"
      ? { target: "pino-pretty", options: { colorize: true } }
      : undefined,
  // Add default fields to every log entry
  base: {
    service: "task-api",
    version: process.env.APP_VERSION || "unknown",
  },
});

export default logger;</code></pre><pre><code>// src/app.ts
import express from "express";
import pinoHttp from "pino-http";
import logger from "./logger";

const app = express();

// Automatically log every HTTP request with method, URL, status, and duration
app.use(pinoHttp({ logger }));

app.get("/tasks", async (req, res) =&gt; {
  try {
    const tasks = await db.query("SELECT * FROM tasks");
    // Info-level log with structured context
    logger.info({ taskCount: tasks.length }, "Tasks retrieved successfully");
    res.json(tasks);
  } catch (error) {
    // Error-level log with the error object and request context
    logger.error(
      { err: error, path: req.path, method: req.method },
      "Failed to retrieve tasks"
    );
    res.status(500).json({ error: "Internal server error" });
  }
});</code></pre><p>With <code>pino-http</code>, every request automatically gets a log entry like this:</p><pre><code>{
  "level": 30,
  "time": 1748870525123,
  "service": "task-api",
  "req": { "method": "GET", "url": "/tasks" },
  "res": { "statusCode": 200 },
  "responseTime": 45,
  "msg": "request completed"
}</code></pre><p>This is exactly the kind of data you can search and filter in a log aggregation system like Loki, Elasticsearch, or CloudWatch Logs. You can query things like &#8220;show me all requests where responseTime &gt; 1000&#8221; or &#8220;show me all error-level logs from the task-api service in the last hour.&#8221;</p><h5><strong>Metrics: counting what matters</strong></h5><p>While logs tell you about individual events, metrics tell you about the overall behavior of your system over time. Metrics are numerical measurements collected at regular intervals.</p><p>There are three core metric types you need to know:</p><blockquote><ul><li><p><strong>Counter</strong>: A value that only goes up. Examples: total number of HTTP requests, total number of errors, total bytes transferred. You usually care about the rate of change (requests per second) rather than the raw value.</p></li><li><p><strong>Gauge</strong>: A value that can go up and down. Examples: current CPU usage, memory usage, number of active connections, queue depth. Gauges represent the current state of something.</p></li><li><p><strong>Histogram</strong>: Measures the distribution of values. Examples: request duration, response size. Histograms let you answer questions like &#8220;what is the 99th percentile latency?&#8221; which is far more useful than the average.</p></li></ul></blockquote><p><strong>Prometheus</strong> is the standard metrics system in the Kubernetes ecosystem. It works with a pull model: instead of your application pushing metrics to a server, Prometheus scrapes your application&#8217;s metrics endpoint at regular intervals (usually every 15 or 30 seconds).</p><p>Here is how the flow works:</p><pre><code>Your App (/metrics endpoint)
  |
  v
Prometheus (scrapes every 15s, stores time-series data)
  |
  v
Grafana (queries Prometheus, renders dashboards)
  |
  v
Alertmanager (receives alerts from Prometheus, sends notifications)</code></pre><p>Let&#8217;s add a <code>/metrics</code> endpoint to our TypeScript API using the <code>prom-client</code> library:</p><pre><code>npm install prom-client</code></pre><pre><code>// src/metrics.ts
import client from "prom-client";

// Create a registry to hold all metrics
const register = new client.Registry();

// Add default Node.js metrics (CPU, memory, event loop lag, etc.)
client.collectDefaultMetrics({ register });

// Custom counter: total HTTP requests, labeled by method, path, and status
export const httpRequestsTotal = new client.Counter({
  name: "http_requests_total",
  help: "Total number of HTTP requests",
  labelNames: ["method", "path", "status"] as const,
  registers: [register],
});

// Custom histogram: request duration in seconds
export const httpRequestDuration = new client.Histogram({
  name: "http_request_duration_seconds",
  help: "Duration of HTTP requests in seconds",
  labelNames: ["method", "path", "status"] as const,
  buckets: [0.01, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10],
  registers: [register],
});

// Custom gauge: number of active database connections
export const dbActiveConnections = new client.Gauge({
  name: "db_active_connections",
  help: "Number of active database connections",
  registers: [register],
});

export { register };</code></pre><pre><code>// src/middleware/metrics.ts
import { Request, Response, NextFunction } from "express";
import { httpRequestsTotal, httpRequestDuration } from "../metrics";

export function metricsMiddleware(
  req: Request,
  res: Response,
  next: NextFunction
) {
  const start = Date.now();

  res.on("finish", () =&gt; {
    const duration = (Date.now() - start) / 1000;
    const path = req.route?.path || req.path;
    const labels = {
      method: req.method,
      path: path,
      status: res.statusCode.toString(),
    };

    httpRequestsTotal.inc(labels);
    httpRequestDuration.observe(labels, duration);
  });

  next();
}</code></pre><pre><code>// src/app.ts - add the metrics endpoint and middleware
import { register } from "./metrics";
import { metricsMiddleware } from "./middleware/metrics";

// Apply metrics middleware to all routes
app.use(metricsMiddleware);

// Expose metrics for Prometheus to scrape
app.get("/metrics", async (_req, res) =&gt; {
  res.set("Content-Type", register.contentType);
  res.end(await register.metrics());
});</code></pre><p>When Prometheus scrapes <code>/metrics</code>, it gets output like this:</p><pre><code># HELP http_requests_total Total number of HTTP requests
# TYPE http_requests_total counter
http_requests_total{method="GET",path="/tasks",status="200"} 1523
http_requests_total{method="POST",path="/tasks",status="201"} 47
http_requests_total{method="GET",path="/tasks",status="500"} 3

# HELP http_request_duration_seconds Duration of HTTP requests in seconds
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{method="GET",path="/tasks",status="200",le="0.05"} 1200
http_request_duration_seconds_bucket{method="GET",path="/tasks",status="200",le="0.1"} 1450
http_request_duration_seconds_bucket{method="GET",path="/tasks",status="200",le="0.25"} 1510
http_request_duration_seconds_bucket{method="GET",path="/tasks",status="200",le="+Inf"} 1523</code></pre><p>For Prometheus to discover this endpoint in Kubernetes, you add annotations to your pod or service:</p><pre><code># In your Helm chart's deployment template or values
metadata:
  annotations:
    prometheus.io/scrape: "true"
    prometheus.io/port: "3000"
    prometheus.io/path: "/metrics"</code></pre><h5><strong>Installing Prometheus and Grafana on EKS</strong></h5><p>The easiest way to get Prometheus and Grafana running on Kubernetes is the <code>kube-prometheus-stack</code> Helm chart. This single chart installs Prometheus, Grafana, Alertmanager, node-exporter (for host metrics), kube-state-metrics (for Kubernetes object metrics), and a bunch of pre-configured dashboards and alerting rules.</p><pre><code># Add the Prometheus community Helm repository
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update

# Create a namespace for monitoring
kubectl create namespace monitoring

# Install the kube-prometheus-stack
helm install monitoring prometheus-community/kube-prometheus-stack \
  --namespace monitoring \
  --set grafana.adminPassword=your-secure-password \
  --set prometheus.prometheusSpec.retention=7d \
  --set prometheus.prometheusSpec.storageSpec.volumeClaimTemplate.spec.resources.requests.storage=20Gi</code></pre><p>That is it. A single Helm command and you have a full monitoring stack. Let&#8217;s verify everything is running:</p><pre><code># Check all pods in the monitoring namespace
kubectl get pods -n monitoring

# Expected output:
# NAME                                                     READY   STATUS    RESTARTS   AGE
# alertmanager-monitoring-kube-prometheus-alertmanager-0    2/2     Running   0          2m
# monitoring-grafana-6c4f8d5b7-x2k4f                      3/3     Running   0          2m
# monitoring-kube-prometheus-operator-7d9f5b8c9-abc12      1/1     Running   0          2m
# monitoring-kube-state-metrics-5f8d9b7c6-def34            1/1     Running   0          2m
# monitoring-prometheus-node-exporter-ghij5                1/1     Running   0          2m
# prometheus-monitoring-kube-prometheus-prometheus-0        2/2     Running   0          2m</code></pre><p>To access Grafana locally, use port-forwarding:</p><pre><code># Forward Grafana to localhost:3001
kubectl port-forward svc/monitoring-grafana 3001:80 -n monitoring

# Open http://localhost:3001 in your browser
# Login: admin / your-secure-password</code></pre><p>For production, you would expose Grafana through an Ingress with TLS. Here is a quick values file for a production-like setup:</p><pre><code># monitoring-values.yaml
grafana:
  adminPassword: "${GRAFANA_ADMIN_PASSWORD}"
  ingress:
    enabled: true
    ingressClassName: alb
    hosts:
      - grafana.yourdomain.com
    tls:
      - secretName: grafana-tls
        hosts:
          - grafana.yourdomain.com

prometheus:
  prometheusSpec:
    retention: 15d
    storageSpec:
      volumeClaimTemplate:
        spec:
          storageClassName: gp3
          resources:
            requests:
              storage: 50Gi
    # Tell Prometheus to scrape pods with the standard annotations
    podMonitorSelectorNilUsesHelmValues: false
    serviceMonitorSelectorNilUsesHelmValues: false

alertmanager:
  alertmanagerSpec:
    storage:
      volumeClaimTemplate:
        spec:
          storageClassName: gp3
          resources:
            requests:
              storage: 5Gi</code></pre><pre><code># Install with the production values
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
  --namespace monitoring \
  -f monitoring-values.yaml</code></pre><h5><strong>PromQL basics: querying your metrics</strong></h5><p>PromQL is the query language for Prometheus. It looks strange at first, but you only need to learn a handful of patterns to cover most use cases.</p><p><strong>Instant vector</strong> - select the current value of a metric:</p><pre><code># All HTTP requests from the task-api
http_requests_total{service="task-api"}

# Only 500 errors
http_requests_total{service="task-api", status="500"}</code></pre><p><strong>Rate</strong> - the most important function. Calculates the per-second rate of increase for counters over a time window:</p><pre><code># Requests per second over the last 5 minutes
rate(http_requests_total[5m])

# Error rate (500s only) per second
rate(http_requests_total{status="500"}[5m])</code></pre><p><strong>Aggregation</strong> - combine multiple time series:</p><pre><code># Total requests per second across all instances
sum(rate(http_requests_total[5m]))

# Requests per second grouped by status code
sum by (status) (rate(http_requests_total[5m]))

# Error percentage
sum(rate(http_requests_total{status=~"5.."}[5m]))
/
sum(rate(http_requests_total[5m]))
* 100</code></pre><p><strong>Histogram quantiles</strong> - calculate percentiles:</p><pre><code># p99 latency (99th percentile)
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))

# p50 latency (median)
histogram_quantile(0.50, rate(http_request_duration_seconds_bucket[5m]))

# p99 latency per endpoint
histogram_quantile(0.99, sum by (path, le) (rate(http_request_duration_seconds_bucket[5m])))</code></pre><p>Here are some queries you will use all the time:</p><pre><code># CPU usage by pod (percentage)
sum by (pod) (rate(container_cpu_usage_seconds_total{namespace="task-api"}[5m])) * 100

# Memory usage by pod (megabytes)
sum by (pod) (container_memory_working_set_bytes{namespace="task-api"}) / 1024 / 1024

# Pod restarts (a restart usually means something crashed)
increase(kube_pod_container_status_restarts_total{namespace="task-api"}[1h])

# Available replicas vs desired replicas (are all pods healthy?)
kube_deployment_status_replicas_available{namespace="task-api"}
/
kube_deployment_spec_replicas{namespace="task-api"}</code></pre><h5><strong>Building a Grafana dashboard</strong></h5><p>Grafana comes with hundreds of pre-built dashboards you can import. For Kubernetes, the kube-prometheus-stack already includes dashboards for node metrics, pod metrics, and cluster overview. But you will also want a custom dashboard for your application.</p><p><strong>Importing a community dashboard:</strong></p><ol><li><p>Open Grafana and go to Dashboards &gt; Import.</p></li><li><p>Enter a dashboard ID from <a href="https://grafana.com/grafana/dashboards/">grafana.com/dashboards</a>. For example, dashboard <code>315</code> is a popular Kubernetes cluster monitoring dashboard.</p></li><li><p>Select your Prometheus data source and click Import.</p></li></ol><p>That gives you a ready-made dashboard in seconds. Now let&#8217;s build a custom one for our API.</p><p><strong>Creating a custom dashboard:</strong></p><ol><li><p>Go to Dashboards &gt; New Dashboard &gt; Add visualization.</p></li><li><p>Select your Prometheus data source.</p></li><li><p>For the first panel, enter this PromQL query:</p></li></ol><pre><code>sum by (status) (rate(http_requests_total{service="task-api"}[5m]))</code></pre><ol start="4"><li><p>Set the panel title to &#8220;Request Rate by Status Code&#8221;.</p></li><li><p>Choose the &#8220;Time series&#8221; visualization type.</p></li><li><p>Under Legend, set it to <code>{{status}}</code> so each line is labeled with its status code.</p></li></ol><p>Add more panels for the metrics that matter most:</p><blockquote><ul><li><p><strong>Request rate</strong>: <code>sum(rate(http_requests_total{service="task-api"}[5m]))</code> as a stat panel showing total RPS.</p></li><li><p><strong>Error rate percentage</strong>: The error percentage query from earlier, displayed as a gauge with thresholds (green &lt; 1%, yellow &lt; 5%, red &gt;= 5%).</p></li><li><p><strong>p99 latency</strong>: <code>histogram_quantile(0.99, sum by (le) (rate(http_request_duration_seconds_bucket{service="task-api"}[5m])))</code> as a time series chart.</p></li><li><p><strong>Active database connections</strong>: <code>db_active_connections{service="task-api"}</code> as a gauge.</p></li><li><p><strong>Pod CPU and memory</strong>: The container queries from the previous section.</p></li></ul></blockquote><p>A good dashboard follows the USE method (Utilization, Saturation, Errors) or the RED method (Rate, Errors, Duration). For an API, the RED method is the most practical:</p><pre><code>RED Dashboard Layout:
+---------------------+-------------------+--------------------+
| Request Rate (RPS)  | Error Rate (%)    | p99 Latency (ms)   |
| [stat panel]        | [gauge panel]     | [stat panel]       |
+---------------------+-------------------+--------------------+
| Request Rate by Status Code (time series)                    |
+--------------------------------------------------------------+
| Latency Distribution: p50, p90, p99 (time series)            |
+--------------------------------------------------------------+
| Error Log Stream (if using Loki)                             |
+--------------------------------------------------------------+</code></pre><p>Once you are happy with the dashboard, save it and note the JSON model. You can export it and store it in your Git repository so it can be provisioned automatically. The kube-prometheus-stack supports dashboard provisioning through ConfigMaps:</p><pre><code># grafana-dashboard-configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: task-api-dashboard
  namespace: monitoring
  labels:
    grafana_dashboard: "1"
data:
  task-api.json: |
    {
      "dashboard": {
        "title": "Task API",
        "panels": [ ... ]
      }
    }</code></pre><h5><strong>Traces: following a request across services</strong></h5><p>Logs tell you what happened in a single service. Traces tell you what happened across multiple services for a single request. Every trace is made up of <strong>spans</strong>, and each span represents a unit of work: an HTTP handler, a database query, an external API call.</p><p>Here is what a trace looks like:</p><pre><code>Trace ID: abc123def456
|
|-- Span: API Gateway (15ms)
|   |-- Span: Authentication middleware (2ms)
|   |-- Span: Orders Service HTTP call (180ms)
|       |-- Span: Database query: SELECT * FROM orders (150ms)  &lt;-- the bottleneck!
|       |-- Span: Cache write (3ms)
|
Total duration: 200ms</code></pre><p>Without tracing, you would see that the API Gateway took 200ms but you would have no idea that the bottleneck was a slow database query inside the Orders Service. With tracing, you can see the exact breakdown.</p><p><strong>OpenTelemetry</strong> (OTel) is the standard for instrumenting applications with traces (and metrics and logs). It provides SDKs for every major language and a vendor-neutral way to export telemetry data. Let&#8217;s add basic tracing to our TypeScript API:</p><pre><code># Install OpenTelemetry packages
npm install @opentelemetry/api \
  @opentelemetry/sdk-node \
  @opentelemetry/auto-instrumentations-node \
  @opentelemetry/exporter-trace-otlp-http</code></pre><pre><code>// src/tracing.ts - must be imported before anything else
import { NodeSDK } from "@opentelemetry/sdk-node";
import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { Resource } from "@opentelemetry/resources";
import {
  ATTR_SERVICE_NAME,
  ATTR_SERVICE_VERSION,
} from "@opentelemetry/semantic-conventions";

const sdk = new NodeSDK({
  resource: new Resource({
    [ATTR_SERVICE_NAME]: "task-api",
    [ATTR_SERVICE_VERSION]: process.env.APP_VERSION || "0.1.0",
  }),
  traceExporter: new OTLPTraceExporter({
    // Send traces to an OTel Collector or Jaeger
    url:
      process.env.OTEL_EXPORTER_OTLP_ENDPOINT ||
      "http://otel-collector:4318/v1/traces",
  }),
  instrumentations: [
    getNodeAutoInstrumentations({
      // Auto-instrument Express, HTTP, and database clients
      "@opentelemetry/instrumentation-express": { enabled: true },
      "@opentelemetry/instrumentation-http": { enabled: true },
      "@opentelemetry/instrumentation-pg": { enabled: true },
    }),
  ],
});

sdk.start();
console.log("OpenTelemetry tracing initialized");

// Graceful shutdown
process.on("SIGTERM", () =&gt; {
  sdk.shutdown().then(() =&gt; process.exit(0));
});</code></pre><pre><code>// src/index.ts - import tracing FIRST
import "./tracing";
import app from "./app";

const port = process.env.PORT || 3000;
app.listen(port, () =&gt; {
  console.log(`Server running on port ${port}`);
});</code></pre><p>With auto-instrumentation, every incoming HTTP request, outgoing HTTP call, and database query automatically gets a span. The SDK propagates the trace context through HTTP headers (<code>traceparent</code>), so when service A calls service B, both services&#8217; spans are linked under the same trace ID.</p><p>For custom spans when you need more detail:</p><pre><code>// src/services/orders.ts
import { trace } from "@opentelemetry/api";

const tracer = trace.getTracer("task-api");

export async function processOrder(orderId: string) {
  // Create a custom span for this operation
  return tracer.startActiveSpan("processOrder", async (span) =&gt; {
    try {
      span.setAttribute("order.id", orderId);

      // Each sub-operation can have its own span
      const order = await tracer.startActiveSpan(
        "fetchOrder",
        async (fetchSpan) =&gt; {
          const result = await db.query("SELECT * FROM orders WHERE id = $1", [
            orderId,
          ]);
          fetchSpan.end();
          return result;
        }
      );

      await tracer.startActiveSpan(
        "validatePayment",
        async (paymentSpan) =&gt; {
          await paymentService.validate(order.paymentId);
          paymentSpan.end();
        }
      );

      span.setAttribute("order.status", "processed");
      span.end();
      return order;
    } catch (error) {
      span.recordException(error as Error);
      span.setStatus({ code: 2, message: (error as Error).message });
      span.end();
      throw error;
    }
  });
}</code></pre><p>To view traces, you need a trace backend. For development, Jaeger is the easiest to set up:</p><pre><code># Run Jaeger locally with Docker
docker run -d --name jaeger \
  -p 16686:16686 \
  -p 4318:4318 \
  jaegertracing/all-in-one:latest

# Open http://localhost:16686 to view traces</code></pre><p>In a Kubernetes cluster, you can deploy Jaeger alongside the OpenTelemetry Collector using the Jaeger Operator or a Helm chart. The kube-prometheus-stack does not include tracing out of the box, but Grafana can connect to Jaeger as a data source and display traces alongside your metrics dashboards.</p><h5><strong>The observability workflow in practice</strong></h5><p>Let&#8217;s walk through a realistic scenario to see how all three pillars work together.</p><p><strong>Scenario</strong>: Users report that creating tasks is slow.</p><p><strong>Step 1: Check the dashboard.</strong> Open your Grafana RED dashboard. You notice that the p99 latency for POST /tasks has jumped from 100ms to 3 seconds in the last 30 minutes. The error rate is still low, so requests are succeeding but they are slow.</p><p><strong>Step 2: Narrow down with metrics.</strong> Add a PromQL query to check if the problem is specific to one pod or all pods:</p><pre><code>histogram_quantile(0.99,
  sum by (pod, le) (
    rate(http_request_duration_seconds_bucket{path="/tasks", method="POST"}[5m])
  )
)</code></pre><p>All pods show the same slow latency, so the issue is not a single unhealthy pod.</p><p><strong>Step 3: Find a slow trace.</strong> Go to Jaeger (or Grafana Tempo) and search for traces where the operation is <code>POST /tasks</code> and the duration is greater than 2 seconds. You find several traces and open one. The trace shows:</p><pre><code>POST /tasks (3.1s)
  |-- Express middleware (2ms)
  |-- insertTask (3.05s)
      |-- pg.query: INSERT INTO tasks... (3.04s)  &lt;-- the problem</code></pre><p>The database INSERT is taking 3 seconds. That is abnormal.</p><p><strong>Step 4: Check the logs.</strong> Search your logs for database-related errors in the last 30 minutes:</p><pre><code>{
  "level": "warn",
  "message": "Slow query detected",
  "query": "INSERT INTO tasks...",
  "duration_ms": 3041,
  "service": "task-api",
  "connection_pool_active": 19,
  "connection_pool_max": 20
}</code></pre><p>The connection pool is almost full. You check further and find that a background job that runs every 30 minutes is holding connections open longer than expected. You fix the background job, and latency returns to normal.</p><p>This is the observability workflow: alert or symptom, dashboard, trace, logs, root cause. Each pillar narrowed the problem until you found the answer.</p><h5><strong>Alerting basics</strong></h5><p>Dashboards are useful for investigation, but you need alerts to know when something is wrong before your users tell you. Prometheus supports alerting rules that evaluate PromQL expressions and fire alerts when conditions are met.</p><p>Here is a PrometheusRule resource for a simple alert:</p><pre><code># alert-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: task-api-alerts
  namespace: monitoring
  labels:
    release: monitoring  # Must match the kube-prometheus-stack release name
spec:
  groups:
    - name: task-api
      rules:
        # Alert when error rate exceeds 5% for 5 minutes
        - alert: HighErrorRate
          expr: |
            sum(rate(http_requests_total{service="task-api", status=~"5.."}[5m]))
            /
            sum(rate(http_requests_total{service="task-api"}[5m]))
            &gt; 0.05
          for: 5m
          labels:
            severity: warning
          annotations:
            summary: "High error rate on task-api"
            description: &gt;
              The task-api error rate is {{ $value | humanizePercentage }}
              over the last 5 minutes.

        # Alert when p99 latency exceeds 1 second for 10 minutes
        - alert: HighLatency
          expr: |
            histogram_quantile(0.99,
              sum by (le) (rate(http_request_duration_seconds_bucket{service="task-api"}[5m]))
            ) &gt; 1
          for: 10m
          labels:
            severity: warning
          annotations:
            summary: "High p99 latency on task-api"
            description: &gt;
              The task-api p99 latency is {{ $value | humanizeDuration }}
              over the last 5 minutes.

        # Alert when a pod has restarted more than 3 times in an hour
        - alert: PodCrashLooping
          expr: |
            increase(kube_pod_container_status_restarts_total{
              namespace="task-api"
            }[1h]) &gt; 3
          for: 5m
          labels:
            severity: critical
          annotations:
            summary: "Pod crash-looping in task-api namespace"
            description: &gt;
              Pod {{ $labels.pod }} has restarted {{ $value }} times
              in the last hour.</code></pre><p>Apply the rule and Prometheus picks it up automatically:</p><pre><code>kubectl apply -f alert-rules.yaml</code></pre><p><strong>Alertmanager</strong> receives alerts from Prometheus and routes them to the right destination: Slack, PagerDuty, email, or a webhook. The kube-prometheus-stack includes Alertmanager. Here is a basic configuration that sends alerts to a Slack channel:</p><pre><code># In your monitoring-values.yaml, add Alertmanager configuration
alertmanager:
  config:
    global:
      slack_api_url: "https://hooks.slack.com/services/YOUR/SLACK/WEBHOOK"
    route:
      receiver: "slack-notifications"
      group_by: ["alertname", "namespace"]
      group_wait: 30s
      group_interval: 5m
      repeat_interval: 4h
    receivers:
      - name: "slack-notifications"
        slack_configs:
          - channel: "#alerts"
            send_resolved: true
            title: '{{ .GroupLabels.alertname }}'
            text: &gt;-
              {{ range .Alerts }}
              *{{ .Annotations.summary }}*
              {{ .Annotations.description }}
              {{ end }}</code></pre><p>The key settings to understand:</p><blockquote><ul><li><p><strong>group_by</strong>: Groups related alerts together so you get one notification instead of fifty when something goes wrong.</p></li><li><p><strong>group_wait</strong>: How long to wait before sending the first notification after a group is created. Gives time for related alerts to arrive and get grouped.</p></li><li><p><strong>repeat_interval</strong>: How often to re-send an unresolved alert. You do not want to get paged every 30 seconds for the same issue.</p></li><li><p><strong>send_resolved</strong>: Sends a notification when the alert clears. Nice to know when the problem is fixed without checking manually.</p></li></ul></blockquote><h5><strong>Connecting the dots: logs, metrics, and traces together</strong></h5><p>The real power of observability comes when you connect all three pillars. The key is the <strong>trace ID</strong>. When a request enters your system, it gets a unique trace ID. If you include that trace ID in your logs and your metrics labels, you can jump from a log entry to the corresponding trace, or from an alert to the exact logs that explain what happened.</p><p>Here is how to add the trace ID to your structured logs:</p><pre><code>// src/middleware/traceContext.ts
import { trace, context } from "@opentelemetry/api";
import { Request, Response, NextFunction } from "express";
import logger from "../logger";

export function traceContextMiddleware(
  req: Request,
  _res: Response,
  next: NextFunction
) {
  const span = trace.getSpan(context.active());
  if (span) {
    const spanContext = span.spanContext();
    // Attach trace ID to the request logger so all logs in this request
    // include the trace ID automatically
    req.log = logger.child({
      traceId: spanContext.traceId,
      spanId: spanContext.spanId,
    });
  }
  next();
}</code></pre><p>Now every log entry from a request includes the trace ID:</p><pre><code>{
  "level": "error",
  "message": "Failed to process order",
  "traceId": "abc123def456789",
  "spanId": "def456789abc123",
  "orderId": "12345",
  "service": "task-api"
}</code></pre><p>In Grafana, you can configure a data link from your log panel (Loki) to your trace panel (Jaeger or Tempo). Click on a log entry and jump directly to the trace. This is the single most useful feature for debugging production issues.</p><h5><strong>What to observe: a starter checklist</strong></h5><p>When you are just getting started, it is easy to get overwhelmed by the number of things you could measure. Here is a practical starting point:</p><blockquote><ul><li><p><strong>For every API endpoint</strong>: Request rate, error rate, and latency (the RED method). These three metrics cover most problems.</p></li><li><p><strong>For your infrastructure</strong>: CPU usage, memory usage, disk usage, and network I/O per pod. The kube-prometheus-stack gives you these for free.</p></li><li><p><strong>For your database</strong>: Active connections, query duration, and connection pool utilization. These are the most common source of application performance issues.</p></li><li><p><strong>For your application health</strong>: Pod restarts, deployment replica status, and container readiness. These tell you if Kubernetes is struggling to keep your app running.</p></li></ul></blockquote><p>Start with these and add more metrics as you encounter specific problems. Do not try to measure everything on day one.</p><h5><strong>Advanced topics</strong></h5><p>We covered the essentials in this article, but observability goes much deeper. Here are topics worth exploring once you are comfortable with the basics:</p><blockquote><ul><li><p><strong>SLO-based alerting</strong>: Instead of alerting on raw thresholds (&#8220;latency &gt; 1s&#8221;), define Service Level Objectives and alert on error budget burn rate. This avoids noisy alerts and focuses on what matters to users.</p></li><li><p><strong>Log aggregation with Loki</strong>: Loki is the logging equivalent of Prometheus. It indexes log metadata (labels) and stores the log content compressed, making it much cheaper than Elasticsearch for Kubernetes logging.</p></li><li><p><strong>Distributed tracing at scale with Tempo</strong>: Grafana Tempo is a trace backend designed to work seamlessly with Grafana, Loki, and Prometheus. It supports trace-to-log and trace-to-metric correlation out of the box.</p></li><li><p><strong>Trace-based testing</strong>: Use traces to verify that your services communicate correctly in integration tests. Tools like Tracetest let you write assertions against trace data.</p></li><li><p><strong>Custom metrics for business logic</strong>: Track things like orders processed, revenue per minute, or user signups. These business metrics are often more valuable than technical metrics.</p></li></ul></blockquote><p>For a comprehensive deep dive into all of these topics, check out the <a href="https://segfault.pw/blog/sre-observability-deep-dive-traces-logs-and-metrics">SRE Observability Deep Dive</a>. It covers OpenTelemetry instrumentation patterns, Loki setup, Grafana Tempo, SLO-based alerting with Pyrra, and production-grade observability architectures.</p><h5><strong>Closing notes</strong></h5><p>Observability is not optional. Once your application is running in production, you need to know what it is doing, how it is performing, and when something goes wrong. The three pillars (logs, metrics, and traces) give you complementary views into your system&#8217;s behavior.</p><p>In this article we covered what observability is and why it matters, the three pillars and when to use each one, structured logging with pino, Prometheus metrics with prom-client, installing Prometheus and Grafana with the kube-prometheus-stack, basic PromQL queries for common scenarios, building Grafana dashboards, distributed tracing with OpenTelemetry, alerting with PrometheusRule and Alertmanager, and the observability workflow for debugging production issues.</p><p>The key takeaway is that observability is a workflow, not a tool. You do not just install Prometheus and call it done. You instrument your application, build dashboards that answer real questions, set up alerts that notify you before your users do, and practice the alert-dashboard-trace-log flow until it becomes second nature.</p><p>In the next article we will cover CI/CD pipelines for Kubernetes, bringing together everything we have built so far into an automated deployment workflow.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: GitOps with ArgoCD]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-gitops-with-argocd</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-gitops-with-argocd</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Sat, 30 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zFxK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zFxK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zFxK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!zFxK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!zFxK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!zFxK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zFxK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043128?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zFxK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!zFxK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!zFxK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!zFxK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0317d1-4241-463d-a581-781ae8a9fa67_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article fourteen of the DevOps from Zero to Hero series. In the previous article we learned how to deploy our TypeScript API to an EKS cluster. We ran <code>kubectl apply</code> and <code>helm install</code> commands to get things running, and that works fine when you are the only person deploying to a single cluster. But what happens when your team grows, when you have multiple environments, or when someone applies a quick fix directly in the cluster and forgets to update the YAML files in Git?</p><p>That is where GitOps comes in. GitOps is a way of managing your Kubernetes deployments where Git is the single source of truth. Instead of running commands against the cluster, you push changes to a Git repository and a controller inside the cluster picks them up and applies them automatically. No more wondering what is running where. No more manual drift. Everything is tracked, reviewed, and reproducible.</p><p>ArgoCD is the most popular GitOps tool for Kubernetes. It is a CNCF graduated project with an excellent web UI, a powerful CLI, and native support for Helm, Kustomize, and plain YAML. In this article we will install ArgoCD on our EKS cluster, deploy our TypeScript API through it, and learn how the whole sync and reconciliation loop works.</p><p>If you are already comfortable with GitOps and want to learn about advanced patterns like ApplicationSets, App of Apps, sync waves, multi-cluster management, RBAC, and notifications, check out <a href="https://segfault.pw/blog/sre-gitops-with-argocd">GitOps with ArgoCD</a> from the SRE series. This article stays beginner-friendly and focuses on getting you from zero to a working GitOps setup.</p><p>Let&#8217;s get into it.</p><h5><strong>What is GitOps?</strong></h5><p>GitOps is an operational model for Kubernetes where you declare what you want running in your cluster in a Git repository, and a controller running inside the cluster continuously makes sure the real state matches the declared state. If someone changes something manually or if a pod crashes and gets recreated with different settings, the controller detects the drift and fixes it.</p><p>This is different from the traditional CI/CD approach where a pipeline runs <code>kubectl apply</code> or <code>helm upgrade</code> at the end of a build. With that push-based model, the CI system needs credentials to your cluster, drift goes undetected, and there is no easy way to know exactly what is running right now. With GitOps, the flow is reversed: the controller lives inside the cluster, pulls the desired state from Git, and handles the apply step itself.</p><pre><code># Traditional push-based CI/CD:
# Developer -&gt; Git push -&gt; CI builds -&gt; CI runs kubectl apply -&gt; Cluster
#                                       (CI needs cluster credentials)
#                                       (manual changes go undetected)

# Pull-based GitOps:
# Developer -&gt; Git push -&gt; Controller detects change -&gt; Controller applies -&gt; Cluster
#                          (controller lives in the cluster)
#                          (drift is detected and corrected automatically)</code></pre><h5><strong>GitOps principles</strong></h5><p>There are four core principles that define a GitOps workflow:</p><blockquote><ul><li><p><strong>Declarative</strong>: Your entire system is described as YAML or JSON files in Git. No imperative scripts, no manual steps, no one-off commands. You declare what you want, not how to get there.</p></li><li><p><strong>Versioned and immutable</strong>: Every change goes through Git, which means every change is versioned, has an author, has a timestamp, and can be reviewed in a pull request. You get a full audit trail for free.</p></li><li><p><strong>Pulled automatically</strong>: A controller running in your cluster watches the Git repository and pulls changes as they appear. You do not push to the cluster. This is more secure because cluster credentials never leave the cluster.</p></li><li><p><strong>Continuously reconciled</strong>: The controller does not apply changes once and forget about them. It runs a loop that constantly compares the live state with the desired state. If they differ for any reason, it corrects the drift.</p></li></ul></blockquote><p>The big win here is that Git becomes the single source of truth. If you want to know what is running in your cluster, look at Git. If you want to roll back, revert a commit. If you want to audit who changed what and when, check the Git history. Everything flows through the same process: commit, push, review, merge, and the controller takes care of the rest.</p><h5><strong>Why ArgoCD</strong></h5><p>There are several GitOps tools out there (Flux is another popular one), but ArgoCD has become the go-to choice for most teams. Here is why:</p><blockquote><ul><li><p><strong>CNCF graduated</strong>: ArgoCD is a graduated project in the Cloud Native Computing Foundation, which means it has passed rigorous security audits and has a large, active community.</p></li><li><p><strong>Great web UI</strong>: ArgoCD ships with a dashboard where you can see every application, its sync status, health status, and the resource tree. This is incredibly helpful for debugging and for giving visibility to the whole team.</p></li><li><p><strong>Kubernetes-native</strong>: ArgoCD uses Custom Resource Definitions (CRDs) to define applications. You manage ArgoCD itself with the same tools you use for everything else in Kubernetes.</p></li><li><p><strong>Multi-format support</strong>: ArgoCD works with plain YAML manifests, Helm charts, Kustomize overlays, Jsonnet, and custom plugins. You do not have to change how you write your manifests.</p></li><li><p><strong>CLI and API</strong>: Beyond the UI, ArgoCD has a full CLI and a gRPC/REST API for automation and scripting.</p></li></ul></blockquote><h5><strong>Installing ArgoCD on EKS</strong></h5><p>We are going to install ArgoCD on the EKS cluster we set up in the previous article. The recommended way is using Helm. First, add the Argo Helm repository and create the namespace:</p><pre><code># Add the ArgoCD Helm repository
helm repo add argo https://argoproj.github.io/argo-helm
helm repo update

# Create the argocd namespace
kubectl create namespace argocd</code></pre><p>Now create a values file to configure the installation. We will keep it simple for now:</p><pre><code># argocd-values.yaml
configs:
  params:
    # If you are terminating TLS at the load balancer or ingress,
    # set this so ArgoCD does not try to handle TLS itself
    server.insecure: true

server:
  service:
    type: LoadBalancer</code></pre><p>This is a minimal configuration. We set <code>server.insecure: true</code> because in a typical EKS setup you terminate TLS at the load balancer or ingress controller level. We also set the service type to LoadBalancer so you can access the UI from your browser.</p><p>Install ArgoCD with Helm:</p><pre><code>helm install argocd argo/argo-cd \
  --namespace argocd \
  --values argocd-values.yaml \
  --wait

# Check that all pods are running
kubectl get pods -n argocd</code></pre><p>You should see something like this:</p><pre><code>NAME                                                READY   STATUS    RESTARTS   AGE
argocd-application-controller-0                     1/1     Running   0          2m
argocd-repo-server-6b7f8d7b4-x9k2l                 1/1     Running   0          2m
argocd-server-7c4f8b6d9-m3n8p                       1/1     Running   0          2m
argocd-redis-5b6c7d8e9-q4r7s                        1/1     Running   0          2m
argocd-applicationset-controller-8f9a1b2c3-t5u6v    1/1     Running   0          2m
argocd-notifications-controller-4d5e6f7a8-w9x0y     1/1     Running   0          2m</code></pre><p>Now get the initial admin password and the load balancer URL:</p><pre><code># Get the initial admin password
kubectl -n argocd get secret argocd-initial-admin-secret \
  -o jsonpath="{.data.password}" | base64 -d
# Save this password, you will need it to log in

# Get the load balancer URL
kubectl -n argocd get svc argocd-server \
  -o jsonpath="{.status.loadBalancer.ingress[0].hostname}"</code></pre><p>Open that URL in your browser and log in with username <code>admin</code> and the password you just retrieved. You should see the ArgoCD dashboard with no applications yet. We will create one shortly.</p><p>You can also install the ArgoCD CLI for managing things from the terminal:</p><pre><code># macOS
brew install argocd

# Linux
curl -sSL -o argocd https://github.com/argoproj/argo-cd/releases/latest/download/argocd-linux-amd64
chmod +x argocd
sudo mv argocd /usr/local/bin/

# Log in to your ArgoCD instance
argocd login &lt;load-balancer-url&gt; --username admin --password &lt;your-password&gt; --insecure</code></pre><p>Once logged in, change the default password:</p><pre><code>argocd account update-password</code></pre><h5><strong>ArgoCD concepts</strong></h5><p>Before we create our first application, let&#8217;s understand the key concepts. ArgoCD has a handful of building blocks that you will use all the time:</p><blockquote><ul><li><p><strong>Application</strong>: The fundamental unit in ArgoCD. An Application defines a source (a Git repository with manifests or a Helm chart), a destination (a Kubernetes cluster and namespace), and a sync policy. Each Application represents one deployable unit.</p></li><li><p><strong>Project</strong>: A logical grouping of Applications with access controls. Projects define which repositories and clusters an Application can use. The <code>default</code> project allows everything, which is fine for getting started.</p></li><li><p><strong>Repository</strong>: A Git repository that ArgoCD watches. You register repositories with ArgoCD so it knows where to pull manifests from. Public repositories work out of the box. Private repositories need credentials.</p></li><li><p><strong>Sync</strong>: The process of applying the desired state from Git to the cluster. When ArgoCD detects a difference between what is in Git and what is running in the cluster, it can sync (apply the changes) either automatically or when you click a button.</p></li><li><p><strong>Health</strong>: ArgoCD understands Kubernetes resource health. A Deployment is healthy when all replicas are available. A Pod is healthy when it is running and ready. A Service is always healthy. ArgoCD shows you the health of every resource in your application.</p></li></ul></blockquote><p>These five concepts cover 90% of what you need to work with ArgoCD day to day. Let&#8217;s put them together by creating our first Application.</p><h5><strong>Setting up a GitOps repository</strong></h5><p>The first thing you need is a Git repository that contains your Kubernetes manifests. This is the repository ArgoCD will watch. You can use the same repository as your application code, but the common practice is to have a separate repository for deployment manifests. This separation makes the workflow cleaner: application code changes trigger CI builds that produce new images, and deployment manifest changes trigger ArgoCD syncs.</p><p>Let&#8217;s create a simple GitOps repository structure:</p><pre><code># Create and initialize the repository
mkdir gitops-repo &amp;&amp; cd gitops-repo
git init
mkdir -p apps/task-api</code></pre><p>Now create the Kubernetes manifests for our TypeScript API. We will use plain YAML to keep things simple, but remember that ArgoCD also supports Helm charts (which we built in article twelve).</p><pre><code># apps/task-api/namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: task-api</code></pre><pre><code># apps/task-api/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: task-api
  namespace: task-api
  labels:
    app: task-api
spec:
  replicas: 2
  selector:
    matchLabels:
      app: task-api
  template:
    metadata:
      labels:
        app: task-api
    spec:
      containers:
        - name: task-api
          image: ghcr.io/your-org/task-api:v1.0.0
          ports:
            - containerPort: 3000
          env:
            - name: NODE_ENV
              value: production
            - name: PORT
              value: "3000"
          resources:
            requests:
              cpu: 100m
              memory: 128Mi
            limits:
              memory: 256Mi
          readinessProbe:
            httpGet:
              path: /health
              port: 3000
            initialDelaySeconds: 5
            periodSeconds: 10
          livenessProbe:
            httpGet:
              path: /health
              port: 3000
            initialDelaySeconds: 15
            periodSeconds: 20</code></pre><pre><code># apps/task-api/service.yaml
apiVersion: v1
kind: Service
metadata:
  name: task-api
  namespace: task-api
spec:
  selector:
    app: task-api
  ports:
    - port: 80
      targetPort: 3000
  type: ClusterIP</code></pre><p>Commit and push these files to your Git repository:</p><pre><code>git add .
git commit -m "Add task-api manifests"
git remote add origin https://github.com/your-org/gitops-repo.git
git push -u origin main</code></pre><h5><strong>Creating your first ArgoCD Application</strong></h5><p>Now let&#8217;s tell ArgoCD about our application. You can do this through the UI, the CLI, or by applying a YAML manifest. We will use the YAML manifest approach because it is declarative, versionable, and follows the GitOps philosophy.</p><pre><code># application.yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: task-api
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/your-org/gitops-repo
    targetRevision: main
    path: apps/task-api
  destination:
    server: https://kubernetes.default.svc
    namespace: task-api
  syncPolicy:
    syncOptions:
      - CreateNamespace=true</code></pre><p>Let&#8217;s break down what each field means:</p><blockquote><ul><li><p><strong>metadata.namespace</strong>: Applications always live in the <code>argocd</code> namespace, regardless of where they deploy resources.</p></li><li><p><strong>spec.project</strong>: We use <code>default</code>, which has no restrictions. In a real team setup you would create dedicated projects with scoped access.</p></li><li><p><strong>spec.source.repoURL</strong>: The Git repository ArgoCD watches.</p></li><li><p><strong>spec.source.targetRevision</strong>: The branch, tag, or commit to track. Using <code>main</code> means ArgoCD follows the tip of the main branch.</p></li><li><p><strong>spec.source.path</strong>: The directory inside the repository that contains the manifests.</p></li><li><p><strong>spec.destination.server</strong>: The Kubernetes API server to deploy to. <code>https://kubernetes.default.svc</code> means the same cluster where ArgoCD is running.</p></li><li><p><strong>spec.destination.namespace</strong>: The target namespace for the deployed resources.</p></li><li><p><strong>syncPolicy.syncOptions</strong>: <code>CreateNamespace=true</code> tells ArgoCD to create the namespace if it does not exist.</p></li></ul></blockquote><p>Apply it:</p><pre><code>kubectl apply -f application.yaml</code></pre><p>If you open the ArgoCD UI now, you will see the <code>task-api</code> application. Its status will be <strong>OutOfSync</strong> because we have not synced it yet. Let&#8217;s do that.</p><h5><strong>The sync loop: how ArgoCD detects drift and reconciles</strong></h5><p>ArgoCD runs a reconciliation loop every three minutes by default. Here is what happens during each cycle:</p><blockquote><ul><li><p><strong>Step 1</strong>: The Application Controller reads the Application CRD and asks the Repository Server to clone the Git repo and render the manifests from the specified path.</p></li><li><p><strong>Step 2</strong>: The Repository Server fetches the latest commit from the branch (or tag), reads the YAML files, and returns the rendered manifests. If you are using Helm, it runs <code>helm template</code>. If you are using Kustomize, it runs <code>kustomize build</code>.</p></li><li><p><strong>Step 3</strong>: The Application Controller compares the rendered manifests with the live state of the resources in the cluster. It does a field-by-field comparison to detect any differences.</p></li><li><p><strong>Step 4</strong>: If there are differences, ArgoCD marks the application as <strong>OutOfSync</strong> and shows you exactly what changed. Depending on your sync policy, it either waits for you to manually trigger a sync or applies the changes automatically.</p></li></ul></blockquote><pre><code>Reconciliation loop (every 3 minutes):

  Git repository          ArgoCD                  Kubernetes cluster
  &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;    &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;        &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
  &#9474; YAML files   &#9474;&#9472;&#9472;&#9472;&gt;&#9474; Repo Server   &#9474;        &#9474; Live resources   &#9474;
  &#9474; (desired     &#9474;    &#9474; (renders      &#9474;        &#9474; (actual state)   &#9474;
  &#9474;  state)      &#9474;    &#9474;  manifests)   &#9474;        &#9474;                  &#9474;
  &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;    &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;        &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                              &#9474;                         &#9474;
                              v                         &#9474;
                      &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;                 &#9474;
                      &#9474; App Controller &#9474;&lt;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                      &#9474; (compares     &#9474;
                      &#9474;  desired vs   &#9474;
                      &#9474;  actual)      &#9474;
                      &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9516;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                              &#9474;
                      OutOfSync? &#9472;&#9472;&gt; Sync (apply changes)
                      Synced?   &#9472;&#9472;&gt; Do nothing</code></pre><p>This continuous loop is what makes GitOps powerful. If someone runs <code>kubectl edit</code> and changes a replica count directly in the cluster, ArgoCD will detect the drift and either alert you or fix it automatically (depending on your configuration).</p><h5><strong>Manual sync vs auto-sync</strong></h5><p>When we created our Application above, we did not enable auto-sync. This means ArgoCD will detect changes but wait for you to manually trigger the sync. Let&#8217;s do our first manual sync:</p><pre><code># Sync using the CLI
argocd app sync task-api

# Or you can click the "Sync" button in the ArgoCD UI</code></pre><p>ArgoCD will apply all the manifests from the Git repository to the cluster. You can watch the progress in the UI or with the CLI:</p><pre><code># Watch the sync progress
argocd app get task-api

# Check that the pods are running
kubectl get pods -n task-api</code></pre><p>After the sync completes, the application status should show <strong>Synced</strong> and <strong>Healthy</strong>. Now let&#8217;s talk about when to use manual sync versus auto-sync.</p><p><strong>Manual sync</strong> is good for:</p><blockquote><ul><li><p><strong>Production environments</strong> where you want a human to review and approve every deployment.</p></li><li><p><strong>Initial setup</strong> when you are getting comfortable with ArgoCD and want to see what it will do before it does it.</p></li><li><p><strong>Sensitive applications</strong> where you need an extra layer of control.</p></li></ul></blockquote><p><strong>Auto-sync</strong> is good for:</p><blockquote><ul><li><p><strong>Development and staging environments</strong> where you want changes to be applied as soon as they are merged to the main branch.</p></li><li><p><strong>Infrastructure components</strong> that should always match what is in Git (monitoring, logging, ingress controllers).</p></li><li><p><strong>Teams that have a solid review process</strong> and trust that anything merged to main is ready to deploy.</p></li></ul></blockquote><p>To enable auto-sync, update the Application manifest:</p><pre><code># application.yaml (with auto-sync enabled)
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: task-api
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/your-org/gitops-repo
    targetRevision: main
    path: apps/task-api
  destination:
    server: https://kubernetes.default.svc
    namespace: task-api
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
    syncOptions:
      - CreateNamespace=true</code></pre><p>The two new fields under <code>automated</code> are important:</p><blockquote><ul><li><p><strong>prune</strong>: When set to <code>true</code>, ArgoCD will delete resources from the cluster that no longer exist in Git. If you remove a ConfigMap from your Git repository, ArgoCD removes it from the cluster too. Without this, deleted resources would linger forever.</p></li><li><p><strong>selfHeal</strong>: When set to <code>true</code>, ArgoCD will revert any manual changes made to the cluster. If someone runs <code>kubectl scale deployment task-api --replicas=5</code> directly, ArgoCD will detect the drift and set it back to whatever is declared in Git.</p></li></ul></blockquote><p>Apply the updated manifest:</p><pre><code>kubectl apply -f application.yaml</code></pre><p>From now on, every time you push a change to the <code>apps/task-api</code> directory in the <code>main</code> branch, ArgoCD will automatically apply it to the cluster within three minutes (or sooner if you configure a webhook).</p><h5><strong>Deploying the TypeScript API with a Helm chart</strong></h5><p>In article twelve we created a Helm chart for our TypeScript API. ArgoCD has native Helm support, so you can point an Application directly at a Helm chart in a Git repository. Let&#8217;s set that up.</p><p>Assuming your GitOps repository has the Helm chart at <code>charts/task-api/</code>, create an Application that uses it:</p><pre><code># application-helm.yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: task-api-helm
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/your-org/gitops-repo
    targetRevision: main
    path: charts/task-api
    helm:
      releaseName: task-api
      valueFiles:
        - values-production.yaml
  destination:
    server: https://kubernetes.default.svc
    namespace: task-api
  syncPolicy:
    automated:
      prune: true
      selfHeal: true
    syncOptions:
      - CreateNamespace=true</code></pre><p>The <code>spec.source.helm</code> section is where the Helm-specific configuration goes. <code>releaseName</code> is the name Helm uses for the release, and <code>valueFiles</code> points to a values file relative to the chart directory. You can also inline values directly:</p><pre><code>    helm:
      releaseName: task-api
      values: |
        replicaCount: 3
        image:
          repository: ghcr.io/your-org/task-api
          tag: v1.2.0
        resources:
          requests:
            cpu: 100m
            memory: 128Mi
          limits:
            memory: 256Mi</code></pre><p>This is how most teams handle deployments in practice: the Helm chart lives in the GitOps repository (or in an OCI registry), and ArgoCD renders and applies it. To deploy a new version, you update the image tag in the values file, commit, push, and ArgoCD takes care of the rest.</p><h5><strong>Navigating the ArgoCD UI</strong></h5><p>The ArgoCD web UI is one of its biggest selling points. Let&#8217;s walk through what you will see.</p><p><strong>Application list view</strong>: This is the main page. You see all your applications as cards, each showing the application name, sync status (Synced, OutOfSync, Unknown), health status (Healthy, Degraded, Progressing, Missing), the target revision, and the last sync time. Green means everything is fine. Yellow means something is progressing. Red means something is wrong.</p><p><strong>Application detail view</strong>: Click on an application to see its resource tree. This is a visual representation of every Kubernetes resource managed by the application. For our task-api, you would see the Deployment, which owns a ReplicaSet, which owns the individual Pods. The Service is shown as a separate node. Each resource shows its health status with a colored icon.</p><p><strong>Resource diff view</strong>: Click on any resource to see its details. The &#8220;Diff&#8221; tab shows you exactly what is different between the desired state (from Git) and the live state (in the cluster). This is extremely helpful for debugging sync issues.</p><p><strong>Sync status bar</strong>: At the top of the detail view, you see the current sync status and a &#8220;Sync&#8221; button. If the application is OutOfSync, you can click Sync to trigger a manual sync. You can also choose to sync specific resources instead of the entire application.</p><p><strong>History and rollback</strong>: The &#8220;History&#8221; tab shows every sync operation with the Git commit that triggered it, the time it happened, and whether it succeeded or failed. You can roll back to any previous sync from here.</p><h5><strong>Rollback: reverting to a previous state</strong></h5><p>Things go wrong. A bad image gets deployed, a configuration change breaks something, or a new version has a bug. With GitOps, you have two ways to roll back.</p><p><strong>The GitOps way (recommended)</strong>: Revert the commit in Git. This is the cleanest approach because it keeps Git as the source of truth and creates an audit trail of the rollback:</p><pre><code># Revert the last commit
git revert HEAD --no-edit
git push

# ArgoCD detects the change and syncs automatically (if auto-sync is enabled)
# Or trigger a manual sync:
argocd app sync task-api</code></pre><p><strong>The ArgoCD way (for emergencies)</strong>: Use the ArgoCD CLI or UI to roll back to a previous sync. This is faster but has a caveat: it does not change Git, so if auto-sync is enabled, ArgoCD will eventually re-sync to the latest Git state and undo your rollback:</p><pre><code># View sync history
argocd app history task-api

# Example output:
# ID  DATE                 REVISION
# 3   2026-05-30 10:15:00  abc1234 (main)
# 2   2026-05-29 14:30:00  def5678 (main)
# 1   2026-05-28 09:00:00  ghi9012 (main)

# Roll back to sync ID 2
argocd app rollback task-api 2</code></pre><p>If you use the ArgoCD rollback, make sure to also disable auto-sync first, or the controller will re-apply the latest Git state and undo your rollback:</p><pre><code># Disable auto-sync before rolling back
argocd app set task-api --sync-policy none

# Roll back
argocd app rollback task-api 2

# Fix the issue in Git, then re-enable auto-sync
argocd app set task-api --sync-policy automated --self-heal --auto-prune</code></pre><p>The key takeaway is that <code>git revert</code> is the preferred way to roll back in a GitOps workflow. It keeps everything consistent and leaves a clear record of what happened and why.</p><h5><strong>A typical GitOps workflow</strong></h5><p>Let&#8217;s put it all together and walk through what a typical deployment looks like end to end:</p><blockquote><ul><li><p><strong>Step 1</strong>: A developer opens a pull request that changes the image tag in the deployment manifest (or the Helm values file) from <code>v1.0.0</code> to <code>v1.1.0</code>.</p></li><li><p><strong>Step 2</strong>: The team reviews the change. Because it is just a YAML diff in a pull request, it is easy to see exactly what will change in the cluster.</p></li><li><p><strong>Step 3</strong>: The pull request is merged to main.</p></li><li><p><strong>Step 4</strong>: ArgoCD detects the new commit within three minutes (or immediately if you have a webhook configured). It compares the new desired state with the live state and finds that the image tag differs.</p></li><li><p><strong>Step 5</strong>: If auto-sync is enabled, ArgoCD applies the change. The Deployment gets updated, Kubernetes performs a rolling update, and the new pods come up with the <code>v1.1.0</code> image.</p></li><li><p><strong>Step 6</strong>: ArgoCD marks the application as Synced and Healthy once all pods are running and passing readiness checks.</p></li><li><p><strong>Step 7</strong>: If something goes wrong, the team reverts the commit in Git and ArgoCD rolls back automatically.</p></li></ul></blockquote><p>This workflow gives you code review for infrastructure changes, a full audit trail in Git, automatic deployment, automatic drift detection, and easy rollback. That is a lot of value for a relatively simple setup.</p><h5><strong>Advanced topics: where to go next</strong></h5><p>Once you are comfortable with the basics covered here, there is a lot more ArgoCD can do. Here is a quick overview of advanced topics:</p><blockquote><ul><li><p><strong>App of Apps pattern</strong>: Instead of creating Application manifests one by one, you create a parent Application that manages child Applications. This lets you bootstrap an entire cluster with a single Application.</p></li><li><p><strong>ApplicationSets</strong>: A way to generate multiple Applications from a single template. Useful for deploying the same application across multiple clusters or environments automatically.</p></li><li><p><strong>Sync waves and hooks</strong>: Control the order in which resources are applied. For example, you can ensure that a database migration Job runs before the Deployment starts.</p></li><li><p><strong>RBAC and SSO</strong>: Restrict who can see and sync which applications. Integrate with your identity provider for single sign-on.</p></li><li><p><strong>Notifications</strong>: Send alerts to Slack, email, or other channels when syncs succeed or fail.</p></li></ul></blockquote><p>All of these topics are covered in depth in <a href="https://segfault.pw/blog/sre-gitops-with-argocd">GitOps with ArgoCD</a> from the SRE series. That article goes into ApplicationSet generators, sync wave annotations, RBAC policies with AppProjects, notification templates, monitoring ArgoCD with Prometheus, and more. Once you have the basics down from this article, that is a great next step.</p><h5><strong>Closing notes</strong></h5><p>GitOps with ArgoCD gives you a deployment workflow that is declarative, versioned, automated, and auditable. Instead of running commands against your cluster and hoping everyone follows the same process, you push changes to Git and let ArgoCD handle the rest. Every change is reviewed in a pull request, tracked in Git history, and automatically applied to the cluster.</p><p>In this article we covered what GitOps is and why it matters, installed ArgoCD on an EKS cluster with Helm, learned the core concepts (Application, Project, Sync, Health), created our first Application pointing at a Git repository, understood the reconciliation loop and how ArgoCD detects drift, compared manual sync and auto-sync and when to use each, deployed our TypeScript API using both plain manifests and a Helm chart, explored the ArgoCD UI, and learned how to roll back safely.</p><p>The next article will cover monitoring and observability, because deploying applications is only half the battle. You also need to know if they are healthy and performing well.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: EKS, Running Kubernetes on AWS]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-eks</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-eks</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!R_gr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R_gr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R_gr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!R_gr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!R_gr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!R_gr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R_gr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043132?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!R_gr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!R_gr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!R_gr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!R_gr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1decdcc-6278-43ba-a5f2-8cea10051731_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article thirteen of the DevOps from Zero to Hero series. In the previous article we packaged our TypeScript API as a Helm chart. Now it is time to give that chart a real home on AWS by provisioning an EKS cluster.</p><p>Amazon Elastic Kubernetes Service (EKS) is AWS&#8217;s managed Kubernetes offering. You get a production-grade control plane that AWS patches, scales, and keeps highly available. You only worry about your workloads and the worker nodes that run them. If you have been following the series, you already know how ECS works from article eight. EKS takes a different approach: instead of a proprietary API, you get standard Kubernetes, which means everything you learned in articles eleven and twelve (Kubernetes fundamentals and Helm) applies directly.</p><p>If you want to see how Kubernetes on AWS was done before EKS became the default, check out <a href="https://segfault.pw/blog/from_zero_to_hero_with_kops_and_aws">From zero to hero with kops and AWS</a>. That article covers kops, a tool that provisions self-managed clusters. EKS has since become the go-to choice for most teams because it removes the burden of managing the control plane yourself.</p><p>In this article we will cover what EKS is, compare it with ECS, provision a full cluster with Terraform, explore node group options, set up IAM Roles for Service Accounts, configure Karpenter for autoscaling, install the AWS Load Balancer Controller, deploy our TypeScript API, and discuss storage and cost considerations. Let&#8217;s get into it.</p><h5><strong>What is EKS?</strong></h5><p>EKS gives you a managed Kubernetes control plane. That means AWS runs the API server, etcd, the scheduler, and the controller manager for you. These components run across multiple availability zones for high availability, and AWS handles upgrades, patches, and backups.</p><p>Your responsibilities are:</p><blockquote><ul><li><p><strong>Worker nodes</strong>: You provision the EC2 instances (or Fargate profiles) where your pods run. AWS offers managed node groups that automate the lifecycle of these instances, but you still decide instance types, sizes, and scaling.</p></li><li><p><strong>Networking</strong>: EKS integrates with your VPC. Pods get IP addresses from your VPC subnets using the VPC CNI plugin, which means they are first-class citizens on the network.</p></li><li><p><strong>Add-ons</strong>: Things like the CoreDNS, kube-proxy, and the VPC CNI are installed by default, but you manage their versions and configuration.</p></li><li><p><strong>Workloads</strong>: Everything you deploy, from Deployments to StatefulSets to CronJobs, is your responsibility.</p></li></ul></blockquote><p>The EKS control plane costs $0.10 per hour (about $73 per month). On top of that you pay for whatever compute you use for worker nodes. This is important to keep in mind when we discuss cost later.</p><h5><strong>EKS vs ECS: when to use each</strong></h5><p>Both EKS and ECS run containers on AWS, but they solve the problem differently. Here is how to think about the choice:</p><blockquote><ul><li><p><strong>EKS</strong> is standard Kubernetes. If your team already knows Kubernetes, if you need portability across clouds, or if you are running complex microservice architectures with custom operators, service meshes, or advanced scheduling, EKS is the right pick. The ecosystem is massive, and nearly every tool in the CNCF landscape works out of the box.</p></li><li><p><strong>ECS</strong> is AWS-native. If your workloads are straightforward, if your team is small and does not want to learn Kubernetes, or if you want tight integration with AWS services without extra controllers, ECS is simpler and cheaper (no control plane fee). The Fargate launch type means you do not manage any infrastructure at all.</p></li></ul></blockquote><p>A practical rule of thumb: if you have fewer than five services and no requirement for multi-cloud, start with ECS. If you have a growing platform team, need the Kubernetes ecosystem, or plan to run on multiple providers, go with EKS.</p><p>For this series we are covering both because real teams encounter both. You already deployed to ECS in article eight. Now you will see how EKS compares hands-on.</p><h5><strong>Prerequisites</strong></h5><p>Before we start, make sure you have the following installed:</p><pre><code># AWS CLI v2
aws --version

# Terraform
terraform --version

# kubectl
kubectl version --client

# Helm
helm version

# eksctl (optional but useful for debugging)
eksctl version</code></pre><p>You also need an AWS account with permissions to create VPCs, EKS clusters, IAM roles, and EC2 instances. If you followed article six (AWS from scratch), you already have this set up.</p><h5><strong>Provisioning the VPC with Terraform</strong></h5><p>EKS clusters live inside a VPC. The VPC needs public subnets (for load balancers) and private subnets (for worker nodes). Let&#8217;s start with the network foundation.</p><p>Create a new Terraform project:</p><pre><code>mkdir -p eks-cluster/terraform
cd eks-cluster/terraform</code></pre><p>First, the provider and backend configuration:</p><pre><code># providers.tf
terraform {
  required_version = "&gt;= 1.5"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~&gt; 5.0"
    }
    helm = {
      source  = "hashicorp/helm"
      version = "~&gt; 2.12"
    }
    kubectl = {
      source  = "alx-v/kubectl"
      version = "~&gt; 2.1"
    }
  }
}

provider "aws" {
  region = var.region
}

provider "helm" {
  kubernetes {
    host                   = module.eks.cluster_endpoint
    cluster_ca_certificate = base64decode(module.eks.cluster_certificate_authority_data)

    exec {
      api_version = "client.authentication.k8s.io/v1beta1"
      command     = "aws"
      args        = ["eks", "get-token", "--cluster-name", module.eks.cluster_name]
    }
  }
}</code></pre><p>Now the variables:</p><pre><code># variables.tf
variable "region" {
  description = "AWS region"
  type        = string
  default     = "us-east-1"
}

variable "cluster_name" {
  description = "Name of the EKS cluster"
  type        = string
  default     = "devops-zero-to-hero"
}

variable "cluster_version" {
  description = "Kubernetes version"
  type        = string
  default     = "1.31"
}

variable "vpc_cidr" {
  description = "CIDR block for the VPC"
  type        = string
  default     = "10.0.0.0/16"
}</code></pre><p>And the VPC using the official AWS module:</p><pre><code># vpc.tf
data "aws_availability_zones" "available" {
  filter {
    name   = "opt-in-status"
    values = ["opt-in-not-required"]
  }
}

locals {
  azs = slice(data.aws_availability_zones.available.names, 0, 3)
}

module "vpc" {
  source  = "terraform-aws-modules/vpc/aws"
  version = "~&gt; 5.0"

  name = "${var.cluster_name}-vpc"
  cidr = var.vpc_cidr

  azs             = local.azs
  private_subnets = [for k, v in local.azs : cidrsubnet(var.vpc_cidr, 4, k)]
  public_subnets  = [for k, v in local.azs : cidrsubnet(var.vpc_cidr, 8, k + 48)]
  intra_subnets   = [for k, v in local.azs : cidrsubnet(var.vpc_cidr, 8, k + 52)]

  enable_nat_gateway = true
  single_nat_gateway = true

  public_subnet_tags = {
    "kubernetes.io/role/elb" = 1
  }

  private_subnet_tags = {
    "kubernetes.io/role/internal-elb" = 1
    "karpenter.sh/discovery"         = var.cluster_name
  }

  tags = {
    Project     = "devops-zero-to-hero"
    Environment = "dev"
  }
}</code></pre><p>A few things to note about the subnet tags:</p><blockquote><ul><li><p><strong><code>kubernetes.io/role/elb</code></strong> on public subnets tells the AWS Load Balancer Controller where to place internet-facing ALBs.</p></li><li><p><strong><code>kubernetes.io/role/internal-elb</code></strong> on private subnets is for internal load balancers.</p></li><li><p><strong><code>karpenter.sh/discovery</code></strong> on private subnets lets Karpenter find subnets to launch nodes in.</p></li></ul></blockquote><p>We use a single NAT gateway to keep costs down for a dev environment. In production you would want one per availability zone for redundancy.</p><h5><strong>Provisioning the EKS cluster</strong></h5><p>Now for the main event. We will use the official EKS Terraform module, which wraps a lot of complexity into a clean interface:</p><pre><code># eks.tf
module "eks" {
  source  = "terraform-aws-modules/eks/aws"
  version = "~&gt; 20.0"

  cluster_name    = var.cluster_name
  cluster_version = var.cluster_version

  # Cluster access
  cluster_endpoint_public_access = true

  # Cluster add-ons
  cluster_addons = {
    coredns                = {}
    eks-pod-identity-agent = {}
    kube-proxy             = {}
    vpc-cni                = {}
  }

  vpc_id     = module.vpc.vpc_id
  subnet_ids = module.vpc.private_subnets

  # Give the Terraform identity admin access to the cluster
  enable_cluster_creator_admin_permissions = true

  # Managed node groups
  eks_managed_node_groups = {
    default = {
      instance_types = ["t3.medium"]

      min_size     = 2
      max_size     = 5
      desired_size = 2

      labels = {
        role = "general"
      }

      tags = {
        "karpenter.sh/discovery" = var.cluster_name
      }
    }
  }

  tags = {
    Project     = "devops-zero-to-hero"
    Environment = "dev"
  }
}</code></pre><p>This creates an EKS cluster with a managed node group of two <code>t3.medium</code> instances. Let&#8217;s break down what is happening:</p><blockquote><ul><li><p><strong><code>cluster_endpoint_public_access</code></strong>: Makes the Kubernetes API reachable from the internet. For production you might restrict this to specific CIDR blocks or use a VPN.</p></li><li><p><strong><code>cluster_addons</code></strong>: These are the essential EKS add-ons. CoreDNS handles service discovery, kube-proxy manages network rules, and vpc-cni gives pods VPC-native IP addresses.</p></li><li><p><strong><code>enable_cluster_creator_admin_permissions</code></strong>: Grants the IAM identity that creates the cluster full admin access. Without this, you can lock yourself out.</p></li><li><p><strong><code>eks_managed_node_groups</code></strong>: We define one node group with auto-scaling between 2 and 5 nodes.</p></li></ul></blockquote><h5><strong>Node groups: understanding your options</strong></h5><p>EKS gives you three ways to run your workloads. Each has trade-offs:</p><blockquote><ul><li><p><strong>Managed node groups</strong>: AWS handles the EC2 instance lifecycle. You pick instance types and sizes, and AWS takes care of provisioning, draining, and updating nodes. This is the default choice for most teams. The example above uses managed node groups.</p></li><li><p><strong>Self-managed node groups</strong>: You create and manage the EC2 instances yourself using Auto Scaling Groups. This gives you full control but more operational overhead. Use this only if you need custom AMIs, GPUs with specific drivers, or unusual instance configurations.</p></li><li><p><strong>Fargate profiles</strong>: AWS runs your pods on serverless compute. No EC2 instances to manage at all. Each pod gets its own isolated micro-VM. This is great for batch jobs or workloads with unpredictable scaling, but it has limitations: no DaemonSets, no persistent volumes backed by EBS, and higher per-pod cost compared to well-utilized EC2 instances.</p></li></ul></blockquote><p>For most workloads, start with managed node groups. If you need more sophisticated scaling (which we will set up shortly), add Karpenter on top.</p><h5><strong>IAM Roles for Service Accounts (IRSA)</strong></h5><p>This is one of the most important EKS concepts to understand. Your pods often need to talk to AWS services: reading from S3, writing to DynamoDB, sending messages to SQS. The old approach was to attach IAM policies to the node&#8217;s instance profile, but that means every pod on that node gets the same permissions. That is a security nightmare.</p><p>IRSA solves this by letting you map a Kubernetes ServiceAccount to a specific IAM role. Only pods using that ServiceAccount get those permissions. Here is how it works under the hood:</p><pre><code>Pod (with ServiceAccount annotation)
  --&gt; Kubernetes mounts a projected token
    --&gt; AWS STS validates the token via OIDC
      --&gt; Pod assumes the IAM role
        --&gt; Pod gets temporary AWS credentials</code></pre><p>EKS creates an OpenID Connect (OIDC) provider for your cluster. When a pod starts, Kubernetes injects a signed JWT token. AWS STS validates this token against the OIDC provider and issues temporary credentials for the mapped IAM role. No long-lived credentials, no shared permissions.</p><p>Here is how to set up IRSA for a pod that needs S3 access:</p><pre><code># irsa.tf

# The OIDC provider is created by the EKS module automatically
# We just need to create the IAM role and policy

module "s3_reader_irsa" {
  source  = "terraform-aws-modules/iam/aws//modules/iam-role-for-service-accounts-eks"
  version = "~&gt; 5.0"

  role_name = "${var.cluster_name}-s3-reader"

  role_policy_arns = {
    policy = aws_iam_policy.s3_read.arn
  }

  oidc_providers = {
    main = {
      provider_arn               = module.eks.oidc_provider_arn
      namespace_service_accounts = ["default:s3-reader"]
    }
  }
}

resource "aws_iam_policy" "s3_read" {
  name        = "${var.cluster_name}-s3-read"
  description = "Allow reading from the application S3 bucket"

  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Effect = "Allow"
        Action = [
          "s3:GetObject",
          "s3:ListBucket"
        ]
        Resource = [
          "arn:aws:s3:::my-app-bucket",
          "arn:aws:s3:::my-app-bucket/*"
        ]
      }
    ]
  })
}</code></pre><p>Then in your Kubernetes manifest (or Helm values), you annotate the ServiceAccount:</p><pre><code>apiVersion: v1
kind: ServiceAccount
metadata:
  name: s3-reader
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/devops-zero-to-hero-s3-reader</code></pre><p>Any pod using this ServiceAccount will automatically receive temporary AWS credentials scoped to that IAM role. This is the right way to handle AWS permissions in EKS.</p><h5><strong>Cluster autoscaler vs Karpenter</strong></h5><p>When your workloads grow, you need more nodes. There are two main options for autoscaling nodes in EKS:</p><blockquote><ul><li><p><strong>Cluster Autoscaler</strong>: The traditional Kubernetes approach. It watches for pods that cannot be scheduled due to insufficient resources, then adds nodes from your existing node groups. It works, but it is limited by your pre-defined node group configurations. If you need a GPU instance but your node group only has <code>t3.medium</code>, you are stuck.</p></li><li><p><strong>Karpenter</strong>: AWS&#8217;s open-source node provisioner. Instead of scaling pre-defined node groups, Karpenter looks at pending pod requirements and provisions the right instance type on the fly. It can mix instance types, use Spot instances, and right-size nodes based on actual workload needs. It is faster, smarter, and more cost-effective.</p></li></ul></blockquote><p>For new clusters, Karpenter is the better choice. Let&#8217;s set it up.</p><h5><strong>Setting up Karpenter with Terraform</strong></h5><p>Karpenter needs IAM permissions to launch EC2 instances and manage their lifecycle. The official Karpenter module for Terraform makes this straightforward:</p><pre><code># karpenter.tf
module "karpenter" {
  source  = "terraform-aws-modules/eks/aws//modules/karpenter"
  version = "~&gt; 20.0"

  cluster_name = module.eks.cluster_name

  # Create the IAM role for the Karpenter controller
  enable_v1_permissions = true

  # Create the node IAM role that Karpenter-provisioned nodes will use
  node_iam_role_additional_policies = {
    AmazonSSMManagedInstanceCore = "arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore"
  }

  tags = {
    Project     = "devops-zero-to-hero"
    Environment = "dev"
  }
}

# Install Karpenter using Helm
resource "helm_release" "karpenter" {
  namespace        = "kube-system"
  name             = "karpenter"
  repository       = "oci://public.ecr.aws/karpenter"
  chart            = "karpenter"
  version          = "1.1.1"
  wait             = false

  values = [
    &lt;&lt;-EOT
    serviceAccount:
      name: ${module.karpenter.service_account}
    settings:
      clusterName: ${module.eks.cluster_name}
      clusterEndpoint: ${module.eks.cluster_endpoint}
      interruptionQueue: ${module.karpenter.queue_name}
    EOT
  ]
}</code></pre><p>After Karpenter is installed, you need to define a <code>NodePool</code> and an <code>EC2NodeClass</code> that tell Karpenter what kind of nodes to provision:</p><pre><code># karpenter-nodepool.yaml
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: default
spec:
  template:
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64"]
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["on-demand", "spot"]
        - key: karpenter.k8s.aws/instance-category
          operator: In
          values: ["c", "m", "r", "t"]
        - key: karpenter.k8s.aws/instance-generation
          operator: Gt
          values: ["4"]
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: default
      expireAfter: 720h
  disruption:
    consolidationPolicy: WhenEmptyOrUnderutilized
    consolidateAfter: 1m
  limits:
    cpu: "100"
    memory: 200Gi
---
apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
  name: default
spec:
  amiSelectorTerms:
    - alias: al2023@latest
  role: "KarpenterNodeRole-devops-zero-to-hero"
  subnetSelectorTerms:
    - tags:
        karpenter.sh/discovery: devops-zero-to-hero
  securityGroupSelectorTerms:
    - tags:
        karpenter.sh/discovery: devops-zero-to-hero
  tags:
    Project: devops-zero-to-hero
    ManagedBy: karpenter</code></pre><p>Apply the Karpenter resources after the cluster is ready:</p><pre><code>kubectl apply -f karpenter-nodepool.yaml</code></pre><p>Here is what is happening in this configuration:</p><blockquote><ul><li><p><strong>NodePool</strong>: Defines constraints for nodes. We allow both on-demand and spot instances, restrict to modern instance families (c, m, r, t with generation &gt; 4), and set resource limits so Karpenter does not spin up unlimited compute.</p></li><li><p><strong><code>expireAfter</code></strong>: Nodes are recycled after 30 days. This ensures they pick up the latest AMIs and security patches.</p></li><li><p><strong><code>consolidationPolicy</code></strong>: Karpenter actively consolidates workloads. If nodes are empty or underutilized, it moves pods around and terminates the excess nodes to save cost.</p></li><li><p><strong>EC2NodeClass</strong>: Defines AWS-specific settings like the AMI, IAM role, and subnet/security group selectors.</p></li></ul></blockquote><p>With Karpenter running, you can scale down your managed node group to just one or two nodes for system workloads, and let Karpenter handle everything else dynamically.</p><h5><strong>AWS Load Balancer Controller</strong></h5><p>By default, Kubernetes services of type <code>LoadBalancer</code> create Classic Load Balancers on AWS. These are outdated. The AWS Load Balancer Controller replaces that behavior with modern ALBs (for HTTP/HTTPS) and NLBs (for TCP/UDP).</p><p>The controller watches for Ingress resources and Service annotations, then creates and configures the corresponding AWS load balancers automatically. Let&#8217;s install it:</p><pre><code># alb-controller.tf
module "lb_controller_irsa" {
  source  = "terraform-aws-modules/iam/aws//modules/iam-role-for-service-accounts-eks"
  version = "~&gt; 5.0"

  role_name                              = "${var.cluster_name}-lb-controller"
  attach_load_balancer_controller_policy = true

  oidc_providers = {
    main = {
      provider_arn               = module.eks.oidc_provider_arn
      namespace_service_accounts = ["kube-system:aws-load-balancer-controller"]
    }
  }
}

resource "helm_release" "aws_lb_controller" {
  namespace  = "kube-system"
  name       = "aws-load-balancer-controller"
  repository = "https://aws.github.io/eks-charts"
  chart      = "aws-load-balancer-controller"
  version    = "1.9.2"

  set {
    name  = "clusterName"
    value = module.eks.cluster_name
  }

  set {
    name  = "serviceAccount.name"
    value = "aws-load-balancer-controller"
  }

  set {
    name  = "serviceAccount.annotations.eks\\.amazonaws\\.com/role-arn"
    value = module.lb_controller_irsa.iam_role_arn
  }

  set {
    name  = "vpcId"
    value = module.vpc.vpc_id
  }
}</code></pre><p>Notice how we use IRSA here. The Load Balancer Controller needs permissions to create ALBs, manage target groups, and read subnet tags. Instead of giving those permissions to the node, we create a dedicated IAM role and bind it to the controller&#8217;s ServiceAccount.</p><p>Once installed, you can create Ingress resources that automatically provision ALBs:</p><pre><code>apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: task-api
  annotations:
    kubernetes.io/ingress.class: alb
    alb.ingress.kubernetes.io/scheme: internet-facing
    alb.ingress.kubernetes.io/target-type: ip
    alb.ingress.kubernetes.io/listen-ports: '[{"HTTPS": 443}]'
    alb.ingress.kubernetes.io/certificate-arn: arn:aws:acm:us-east-1:123456789012:certificate/abc-123
spec:
  rules:
    - host: api.example.com
      http:
        paths:
          - path: /
            pathType: Prefix
            backend:
              service:
                name: task-api
                port:
                  number: 3000</code></pre><p>The controller reads the annotations, creates an ALB in your public subnets, attaches the ACM certificate for TLS, and routes traffic to your pods. You do not need to manage load balancers manually anymore.</p><h5><strong>Configuring kubeconfig</strong></h5><p>After the cluster is provisioned, you need to configure kubectl to talk to it. The AWS CLI makes this simple:</p><pre><code># Update your kubeconfig
aws eks update-kubeconfig --region us-east-1 --name devops-zero-to-hero

# Verify the connection
kubectl get nodes</code></pre><p>You should see your managed node group instances:</p><pre><code>NAME                             STATUS   ROLES    AGE   VERSION
ip-10-0-1-42.ec2.internal       Ready    &lt;none&gt;   5m    v1.31.2-eks-7f9249a
ip-10-0-2-87.ec2.internal       Ready    &lt;none&gt;   5m    v1.31.2-eks-7f9249a</code></pre><p>If you work with multiple clusters, you can switch between them using contexts:</p><pre><code># List all contexts
kubectl config get-contexts

# Switch to a specific context
kubectl config use-context arn:aws:eks:us-east-1:123456789012:cluster/devops-zero-to-hero

# Rename a context for convenience
kubectl config rename-context \
  arn:aws:eks:us-east-1:123456789012:cluster/devops-zero-to-hero \
  eks-dev</code></pre><h5><strong>Deploying the TypeScript API to EKS</strong></h5><p>Remember the Helm chart we built in article twelve? Now we put it to use. If you have your chart in an OCI registry, the deployment is a single command:</p><pre><code># Create a namespace for the application
kubectl create namespace task-api

# Install the chart
helm install task-api oci://ghcr.io/your-org/charts/task-api \
  --version 0.1.0 \
  --namespace task-api \
  -f values-eks.yaml</code></pre><p>Here is what the EKS-specific values file looks like:</p><pre><code># values-eks.yaml
replicaCount: 2

image:
  repository: 123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api
  tag: "1.0.0"

service:
  type: ClusterIP
  port: 3000

ingress:
  enabled: true
  className: alb
  annotations:
    alb.ingress.kubernetes.io/scheme: internet-facing
    alb.ingress.kubernetes.io/target-type: ip
    alb.ingress.kubernetes.io/listen-ports: '[{"HTTPS": 443}]'
    alb.ingress.kubernetes.io/certificate-arn: arn:aws:acm:us-east-1:123456789012:certificate/abc-123
    alb.ingress.kubernetes.io/healthcheck-path: /health
  hosts:
    - host: api.example.com
      paths:
        - path: /
          pathType: Prefix

resources:
  requests:
    cpu: 100m
    memory: 128Mi
  limits:
    cpu: 500m
    memory: 256Mi

autoscaling:
  enabled: true
  minReplicas: 2
  maxReplicas: 10
  targetCPUUtilizationPercentage: 70

serviceAccount:
  create: true
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/task-api-role</code></pre><p>After the deployment completes, you can check everything is running:</p><pre><code># Check the pods
kubectl get pods -n task-api
NAME                        READY   STATUS    RESTARTS   AGE
task-api-6d8f9c7b4a-k2m5n   1/1     Running   0          2m
task-api-6d8f9c7b4a-x9p3r   1/1     Running   0          2m

# Check the ingress (the ALB takes a minute or two to provision)
kubectl get ingress -n task-api
NAME       CLASS   HOSTS              ADDRESS                                      PORTS   AGE
task-api   alb     api.example.com    k8s-taskapi-xxxxx.us-east-1.elb.amazonaws.com   80      3m

# Test the endpoint
curl https://api.example.com/health
{"status": "ok"}</code></pre><p>The AWS Load Balancer Controller sees the Ingress resource, creates an ALB, configures target groups pointing to your pod IPs, and attaches the TLS certificate. Traffic flows from the internet through the ALB directly to your pods.</p><h5><strong>Storage: EBS CSI driver</strong></h5><p>If your workloads need persistent storage (databases, caches, file uploads), you need the EBS CSI driver. This driver allows Kubernetes PersistentVolumes to be backed by EBS volumes.</p><p>Add it as an EKS add-on in your Terraform:</p><pre><code># Add to the cluster_addons in eks.tf
cluster_addons = {
  coredns                = {}
  eks-pod-identity-agent = {}
  kube-proxy             = {}
  vpc-cni                = {}
  aws-ebs-csi-driver = {
    service_account_role_arn = module.ebs_csi_irsa.iam_role_arn
  }
}</code></pre><pre><code># ebs-csi.tf
module "ebs_csi_irsa" {
  source  = "terraform-aws-modules/iam/aws//modules/iam-role-for-service-accounts-eks"
  version = "~&gt; 5.0"

  role_name             = "${var.cluster_name}-ebs-csi"
  attach_ebs_csi_policy = true

  oidc_providers = {
    main = {
      provider_arn               = module.eks.oidc_provider_arn
      namespace_service_accounts = ["kube-system:ebs-csi-controller-sa"]
    }
  }
}</code></pre><p>Then create a StorageClass and use it in your workloads:</p><pre><code>apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: gp3
  annotations:
    storageclass.kubernetes.io/is-default-class: "true"
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
  fsType: ext4
volumeBindingMode: WaitForFirstConsumer
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: data-volume
spec:
  accessModes:
    - ReadWriteOnce
  storageClassName: gp3
  resources:
    requests:
      storage: 10Gi</code></pre><p>The <code>WaitForFirstConsumer</code> binding mode is important. It delays volume creation until a pod actually needs it, ensuring the volume is created in the same availability zone as the pod. Without this, you can end up with a volume in one AZ and a pod that needs to run in another.</p><h5><strong>Cost considerations</strong></h5><p>EKS is not cheap, especially compared to ECS with Fargate for small workloads. Here is what you are paying for:</p><blockquote><ul><li><p><strong>Control plane</strong>: $0.10/hour ($73/month). This is fixed regardless of how many nodes you run.</p></li><li><p><strong>Worker nodes</strong>: Standard EC2 pricing. A <code>t3.medium</code> (2 vCPU, 4 GB) runs about $30/month on-demand.</p></li><li><p><strong>Spot instances</strong>: Up to 90% cheaper than on-demand, but can be interrupted. Karpenter makes using Spot easy by diversifying across instance types. Great for stateless workloads, not recommended for databases.</p></li><li><p><strong>NAT gateway</strong>: $32/month plus data transfer. This is often the sneaky cost that surprises people. Use a single NAT gateway for dev, one per AZ for production.</p></li><li><p><strong>Load balancers</strong>: ALBs cost about $16/month plus data transfer. Each Ingress resource can share a single ALB using IngressGroups to avoid provisioning one per service.</p></li><li><p><strong>Data transfer</strong>: Inter-AZ traffic costs $0.01/GB each way. Cross-AZ pod-to-pod communication adds up in chatty microservice architectures.</p></li></ul></blockquote><p>Cost saving tips:</p><blockquote><ul><li><p><strong>Use Karpenter with Spot instances</strong> for stateless workloads. Diversify across many instance types to reduce interruption rates.</p></li><li><p><strong>Right-size your nodes</strong>. Karpenter helps here by picking the optimal instance type for your workload mix.</p></li><li><p><strong>Consolidate ALBs</strong> using IngressGroup annotations so multiple services share one ALB.</p></li><li><p><strong>Use a single NAT gateway</strong> for non-production environments.</p></li><li><p><strong>Set resource requests and limits</strong> on every pod so Karpenter can bin-pack efficiently.</p></li><li><p><strong>Consider Savings Plans or Reserved Instances</strong> for baseline capacity you know you will always need.</p></li></ul></blockquote><p>A minimal EKS dev environment (control plane + 2 <code>t3.medium</code> nodes + NAT gateway + ALB) costs roughly $180/month. A production setup with more nodes, multi-AZ NAT, and monitoring will be significantly more. Compare this to ECS with Fargate where you only pay for the compute your containers actually use.</p><h5><strong>Putting it all together</strong></h5><p>Let&#8217;s run through the full provisioning flow:</p><pre><code># Initialize Terraform
cd eks-cluster/terraform
terraform init

# Review the plan
terraform plan -out=tfplan

# Apply (this takes 15-20 minutes, mostly the EKS cluster creation)
terraform apply tfplan

# Configure kubectl
aws eks update-kubeconfig --region us-east-1 --name devops-zero-to-hero

# Verify the cluster
kubectl get nodes
kubectl get pods -n kube-system

# Apply Karpenter resources
kubectl apply -f karpenter-nodepool.yaml

# Deploy the application
kubectl create namespace task-api
helm install task-api oci://ghcr.io/your-org/charts/task-api \
  --version 0.1.0 \
  --namespace task-api \
  -f values-eks.yaml

# Check everything is running
kubectl get all -n task-api</code></pre><p>After about 20 minutes, you will have a fully functional EKS cluster with managed node groups, Karpenter for dynamic scaling, the AWS Load Balancer Controller for automated ALB provisioning, IRSA for secure pod-level AWS permissions, and the EBS CSI driver for persistent storage.</p><h5><strong>Cleaning up</strong></h5><p>If you are following along and do not want to keep the cluster running, tear it down:</p><pre><code># Remove application resources first
helm uninstall task-api -n task-api
kubectl delete -f karpenter-nodepool.yaml

# Destroy everything with Terraform
terraform destroy</code></pre><p>Always remove Kubernetes resources before destroying the infrastructure. If you destroy the VPC while ALBs still exist, Terraform will hang waiting for the load balancers to be deleted, and you will have to clean them up manually in the AWS console.</p><h5><strong>Closing notes</strong></h5><p>EKS gives you the full power of Kubernetes without the operational burden of managing the control plane. In this article we provisioned a complete cluster with Terraform, configured managed node groups for baseline compute, set up Karpenter for intelligent autoscaling, used IRSA for secure pod-level AWS permissions, installed the AWS Load Balancer Controller for automated ALB management, and deployed our TypeScript API from the Helm chart we built in the previous article.</p><p>The trade-off compared to ECS is complexity and cost. EKS requires more infrastructure knowledge, more moving parts, and a baseline cost even when nothing is running. But in return you get the entire Kubernetes ecosystem, portability across clouds, and the ability to handle complex workloads that would be difficult to model in ECS.</p><p>In the next article we will dive into monitoring and observability, because having a running cluster is only the beginning. You need to know what is happening inside it.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Helm Charts]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-helm-charts</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-helm-charts</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Sun, 24 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Msc0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Msc0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Msc0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!Msc0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!Msc0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!Msc0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Msc0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043134?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Msc0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!Msc0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!Msc0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!Msc0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f9e1b54-52b5-4f9a-91b2-92762df2261b_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article twelve of the DevOps from Zero to Hero series. In the previous articles we learned how to deploy containers to Kubernetes using raw YAML manifests. That works fine when you have a single service with a handful of files, but as your applications grow and you start managing multiple environments, raw manifests become painful to maintain.</p><p>Imagine you have a deployment, a service, an ingress, a configmap, and an HPA. Now multiply that by three environments (dev, staging, production) where only a few values change: the image tag, the replica count, the domain name. Suddenly you are copy-pasting YAML files, searching and replacing values by hand, and praying you did not miss something. This is exactly the problem Helm solves.</p><p>Helm is the package manager for Kubernetes. It lets you define your application as a reusable template, parameterize the parts that change, version the whole thing, and install or upgrade it with a single command. If you have ever used <code>apt</code>, <code>brew</code>, or <code>npm</code>, Helm fills the same role for Kubernetes.</p><p>In this article we will cover what Helm is and why it exists, create a chart from scratch, dive into template syntax and helpers, package a TypeScript API with all the resources it needs, manage releases with install, upgrade, and rollback, push charts to OCI registries, and test our work. If you want to see how Helm was used years ago, check out <a href="https://segfault.pw/blog/getting_started_with_helm">Getting started with Helm</a> and <a href="https://segfault.pw/blog/deploying_my_apps_with_helm">Deploying my apps with Helm</a>, but keep in mind those articles are from 2018 and cover Helm 2 which is now deprecated.</p><p>Let&#8217;s get into it.</p><h5><strong>What is Helm?</strong></h5><p>Helm calls itself the package manager for Kubernetes, and that is a good description. There are four core concepts you need to understand:</p><blockquote><ul><li><p><strong>Chart</strong>: A collection of files that describe a related set of Kubernetes resources. Think of it as a package. A chart contains templates, default values, metadata, and optionally sub-charts for dependencies.</p></li><li><p><strong>Release</strong>: A specific instance of a chart running in your cluster. You can install the same chart multiple times with different configurations, and each installation is a separate release with its own name and history.</p></li><li><p><strong>Repository</strong>: A place where charts are stored and shared. This can be a traditional HTTP server, a Helm repository, or an OCI-compliant container registry (the modern approach).</p></li><li><p><strong>Values</strong>: The configuration parameters that customize a chart for a specific deployment. You provide values to override the chart&#8217;s defaults, and Helm uses them to render the templates into valid Kubernetes manifests.</p></li></ul></blockquote><p>Here is how these pieces fit together:</p><pre><code>Chart (package definition)
  + Values (your configuration)
  = Release (running instance in your cluster)
    &#9500;&#9472;&#9472; Deployment (rendered from template)
    &#9500;&#9472;&#9472; Service (rendered from template)
    &#9500;&#9472;&#9472; Ingress (rendered from template)
    &#9492;&#9472;&#9472; ConfigMap (rendered from template)</code></pre><h5><strong>Why Helm over raw manifests</strong></h5><p>You might be wondering if you really need another tool. Here is what Helm gives you that raw YAML does not:</p><blockquote><ul><li><p><strong>Templating</strong>: Write your manifests once with placeholders, and render them with different values for each environment. No more copy-pasting YAML files.</p></li><li><p><strong>Versioning</strong>: Every chart has a version. Every release tracks which version was installed and what values were used. You always know what is running.</p></li><li><p><strong>Rollback</strong>: Made a bad deployment? <code>helm rollback</code> takes you back to the previous working state in seconds. Helm keeps a history of every release revision.</p></li><li><p><strong>Dependency management</strong>: Your application depends on Redis? Add it as a chart dependency and Helm installs both together.</p></li><li><p><strong>Sharing</strong>: Package your chart and push it to a registry. Anyone on your team (or the world) can install it with a single command.</p></li><li><p><strong>Lifecycle hooks</strong>: Run jobs before or after install, upgrade, or delete. Great for database migrations, cache warming, or health checks.</p></li></ul></blockquote><p>The alternative is managing raw YAML with Kustomize or hand-rolled scripts. Kustomize is built into kubectl and works well for simple overlay scenarios, but it does not give you versioning, rollback, or a release history. For most teams, Helm is the better choice once you move beyond trivial deployments.</p><h5><strong>Helm 3 vs Helm 2: a brief history</strong></h5><p>If you have seen older Helm tutorials, they mention something called Tiller. That was a server-side component that Helm 2 required to run inside your cluster. Tiller had cluster-admin permissions and was a significant security concern.</p><p>Helm 3 (released in November 2019) removed Tiller entirely. Here is what changed:</p><blockquote><ul><li><p><strong>No Tiller</strong>: Helm now talks directly to the Kubernetes API using your kubeconfig credentials. No more deploying a privileged pod into your cluster.</p></li><li><p><strong>Three-way strategic merge</strong>: Helm 3 compares the old manifest, the new manifest, and the live state in the cluster. This means manual changes to resources are detected and handled properly during upgrades.</p></li><li><p><strong>Release namespaces</strong>: Releases are stored as Kubernetes secrets in the namespace where they are deployed, not in a central Tiller namespace.</p></li><li><p><strong>JSON Schema validation</strong>: Charts can include a <code>values.schema.json</code> file to validate user-provided values before rendering.</p></li><li><p><strong>OCI registry support</strong>: Charts can be stored in container registries like Docker Hub, GHCR, or ECR, just like container images.</p></li></ul></blockquote><p>If you are starting today, you will only ever use Helm 3. The <code>helm</code> binary you install from the official site is Helm 3. Helm 2 reached end of life in November 2020, so there is no reason to use it for new projects.</p><h5><strong>Installing Helm</strong></h5><p>Installing Helm is straightforward. Pick the method that matches your system:</p><pre><code># macOS with Homebrew
brew install helm

# Linux with the official install script
curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash

# Arch Linux
pacman -S helm

# Verify the installation
helm version</code></pre><p>You should see output like:</p><pre><code>version.BuildInfo{Version:"v3.17.x", GitCommit:"...", GitTreeState:"clean", GoVersion:"go1.23.x"}</code></pre><h5><strong>Creating a chart from scratch</strong></h5><p>Let&#8217;s create our first chart. Helm provides a scaffolding command:</p><pre><code>helm create task-api</code></pre><p>This creates a directory structure with everything you need:</p><pre><code>task-api/
&#9500;&#9472;&#9472; Chart.yaml          # Metadata: name, version, description
&#9500;&#9472;&#9472; values.yaml         # Default configuration values
&#9500;&#9472;&#9472; charts/             # Sub-charts (dependencies)
&#9500;&#9472;&#9472; templates/          # Kubernetes manifest templates
&#9474;   &#9500;&#9472;&#9472; _helpers.tpl    # Named template helpers
&#9474;   &#9500;&#9472;&#9472; deployment.yaml
&#9474;   &#9500;&#9472;&#9472; service.yaml
&#9474;   &#9500;&#9472;&#9472; ingress.yaml
&#9474;   &#9500;&#9472;&#9472; hpa.yaml
&#9474;   &#9500;&#9472;&#9472; serviceaccount.yaml
&#9474;   &#9500;&#9472;&#9472; NOTES.txt       # Post-install instructions shown to the user
&#9474;   &#9492;&#9472;&#9472; tests/
&#9474;       &#9492;&#9472;&#9472; test-connection.yaml
&#9492;&#9472;&#9472; .helmignore         # Files to exclude when packaging</code></pre><p>Let&#8217;s go through the key files one by one.</p><h5><strong>Chart.yaml: the chart metadata</strong></h5><p>This file defines who your chart is. Think of it like a <code>package.json</code> for Helm:</p><pre><code>apiVersion: v2
name: task-api
description: A Helm chart for the Task API TypeScript application
type: application
version: 0.1.0
appVersion: "1.0.0"</code></pre><blockquote><ul><li><p><strong>apiVersion</strong>: Always <code>v2</code> for Helm 3 charts. Helm 2 used <code>v1</code>.</p></li><li><p><strong>name</strong>: The name of the chart. Must be lowercase and may contain hyphens.</p></li><li><p><strong>description</strong>: A short description displayed when searching repositories.</p></li><li><p><strong>type</strong>: Either <code>application</code> (deploys resources) or <code>library</code> (only provides helpers for other charts).</p></li><li><p><strong>version</strong>: The chart version. Follows semantic versioning. Bump this every time you change the chart.</p></li><li><p><strong>appVersion</strong>: The version of the application being deployed. This is informational and does not affect chart behavior.</p></li></ul></blockquote><p>The distinction between <code>version</code> and <code>appVersion</code> is important. The chart version tracks changes to the chart itself (templates, defaults). The app version tracks which version of your application the chart deploys. They evolve independently.</p><h5><strong>values.yaml: the default configuration</strong></h5><p>This is the most important file in a chart. It defines every configurable parameter with sensible defaults:</p><pre><code>replicaCount: 2

image:
  repository: ghcr.io/your-org/task-api
  pullPolicy: IfNotPresent
  tag: ""

nameOverride: ""
fullnameOverride: ""

serviceAccount:
  create: true
  annotations: {}
  name: ""

service:
  type: ClusterIP
  port: 3000

ingress:
  enabled: false
  className: "traefik"
  annotations: {}
  hosts:
    - host: task-api.example.com
      paths:
        - path: /
          pathType: Prefix
  tls: []

resources:
  limits:
    cpu: 200m
    memory: 256Mi
  requests:
    cpu: 100m
    memory: 128Mi

autoscaling:
  enabled: false
  minReplicas: 2
  maxReplicas: 10
  targetCPUUtilizationPercentage: 80

env:
  NODE_ENV: production
  PORT: "3000"</code></pre><p>A few principles for structuring values:</p><blockquote><ul><li><p><strong>Group related settings</strong>: Put all image-related values under <code>image</code>, all ingress settings under <code>ingress</code>, and so on. This makes it easy to find and override things.</p></li><li><p><strong>Provide sensible defaults</strong>: The chart should work with zero overrides for a basic deployment. Production-specific settings (domain names, resource limits) are what users override.</p></li><li><p><strong>Use flat keys where possible</strong>: Deeply nested values are harder to override with <code>--set</code>. Keep the nesting reasonable.</p></li><li><p><strong>Document with comments</strong>: Add comments explaining what each value does, what valid options are, and what the default means.</p></li></ul></blockquote><h5><strong>Template syntax: Go templates</strong></h5><p>Helm templates use Go&#8217;s <code>text/template</code> package with some extra functions from the Sprig library. If you have never seen Go templates before, here is a quick introduction.</p><p>The basic syntax uses double curly braces <code>{{ }}</code> to insert dynamic content. Everything outside the braces is rendered as-is. Let&#8217;s look at the core patterns:</p><pre><code># Simple value substitution
apiVersion: apps/v1
kind: Deployment
metadata:
  name: {{ .Release.Name }}-api
  labels:
    app: {{ .Chart.Name }}
    version: {{ .Chart.AppVersion }}
spec:
  replicas: {{ .Values.replicaCount }}</code></pre><p>Helm provides several built-in objects that you can access in templates:</p><blockquote><ul><li><p><strong><code>.Values</code></strong>: The merged result of values.yaml and any overrides the user provided. This is where most of your dynamic data comes from.</p></li><li><p><strong><code>.Release</code></strong>: Information about the current release. <code>.Release.Name</code> is the release name, <code>.Release.Namespace</code> is the namespace, <code>.Release.IsUpgrade</code> tells you if this is an upgrade.</p></li><li><p><strong><code>.Chart</code></strong>: Contents of Chart.yaml. <code>.Chart.Name</code>, <code>.Chart.Version</code>, <code>.Chart.AppVersion</code>.</p></li><li><p><strong><code>.Template</code></strong>: Information about the current template file. Mostly used for debugging.</p></li><li><p><strong><code>.Capabilities</code></strong>: Information about the Kubernetes cluster. <code>.Capabilities.APIVersions</code> lets you check if a specific API version exists.</p></li></ul></blockquote><h5><strong>Conditionals and loops</strong></h5><p>Templates support control flow. Here is how to conditionally include an ingress and loop over hosts:</p><pre><code>{{- if .Values.ingress.enabled }}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: {{ include "task-api.fullname" . }}
  {{- with .Values.ingress.annotations }}
  annotations:
    {{- toYaml . | nindent 4 }}
  {{- end }}
spec:
  ingressClassName: {{ .Values.ingress.className }}
  rules:
    {{- range .Values.ingress.hosts }}
    - host: {{ .host | quote }}
      http:
        paths:
          {{- range .paths }}
          - path: {{ .path }}
            pathType: {{ .pathType }}
            backend:
              service:
                name: {{ include "task-api.fullname" $ }}
                port:
                  number: {{ $.Values.service.port }}
          {{- end }}
    {{- end }}
  {{- if .Values.ingress.tls }}
  tls:
    {{- toYaml .Values.ingress.tls | nindent 4 }}
  {{- end }}
{{- end }}</code></pre><p>A few things to notice:</p><blockquote><ul><li><p><strong><code>{{- ... }}</code></strong>: The dash trims whitespace before the tag. Without it, you get blank lines in the output.</p></li><li><p><strong><code>range</code></strong>: Loops over a list or map. Inside the loop, <code>.</code> refers to the current item.</p></li><li><p><strong><code>$</code></strong>: Refers to the root scope. When you are inside a <code>range</code> block, <code>.</code> changes to the current item. Use <code>$</code> to access <code>.Values</code> or <code>.Release</code> from within a loop.</p></li><li><p><strong><code>with</code></strong>: Sets the scope of <code>.</code> to the specified object. If the object is empty, the block is skipped entirely. It works like a combined &#8220;if not empty&#8221; and &#8220;set scope.&#8221;</p></li><li><p><strong><code>toYaml</code></strong>: Converts a Go data structure to YAML. Combined with <code>nindent</code>, it handles indentation correctly.</p></li><li><p><strong><code>quote</code></strong>: Wraps the value in double quotes. Always quote hostnames and strings that might contain special characters.</p></li></ul></blockquote><h5><strong>Helpers: _helpers.tpl</strong></h5><p>The <code>_helpers.tpl</code> file (the underscore prefix tells Helm not to render it as a manifest) contains reusable named templates. These are like functions you can call from any template:</p><pre><code>{{/*
Expand the name of the chart.
*/}}
{{- define "task-api.name" -}}
{{- default .Chart.Name .Values.nameOverride | trunc 63 | trimSuffix "-" }}
{{- end }}

{{/*
Create a default fully qualified app name.
*/}}
{{- define "task-api.fullname" -}}
{{- if .Values.fullnameOverride }}
{{- .Values.fullnameOverride | trunc 63 | trimSuffix "-" }}
{{- else }}
{{- $name := default .Chart.Name .Values.nameOverride }}
{{- if contains $name .Release.Name }}
{{- .Release.Name | trunc 63 | trimSuffix "-" }}
{{- else }}
{{- printf "%s-%s" .Release.Name $name | trunc 63 | trimSuffix "-" }}
{{- end }}
{{- end }}
{{- end }}

{{/*
Common labels
*/}}
{{- define "task-api.labels" -}}
helm.sh/chart: {{ include "task-api.chart" . }}
{{ include "task-api.selectorLabels" . }}
{{- if .Chart.AppVersion }}
app.kubernetes.io/version: {{ .Chart.AppVersion | quote }}
{{- end }}
app.kubernetes.io/managed-by: {{ .Release.Service }}
{{- end }}

{{/*
Selector labels
*/}}
{{- define "task-api.selectorLabels" -}}
app.kubernetes.io/name: {{ include "task-api.name" . }}
app.kubernetes.io/instance: {{ .Release.Name }}
{{- end }}</code></pre><p>You call these named templates using <code>include</code>:</p><pre><code>metadata:
  name: {{ include "task-api.fullname" . }}
  labels:
    {{- include "task-api.labels" . | nindent 4 }}</code></pre><p>Why <code>include</code> instead of <code>template</code>? The <code>template</code> action outputs text directly and cannot be piped to other functions. <code>include</code> returns the output as a string, so you can pipe it to <code>nindent</code>, <code>trim</code>, or any other function. Always prefer <code>include</code> over <code>template</code>.</p><p>The <code>trunc 63</code> calls throughout the helpers are not arbitrary. Kubernetes labels and names have a 63-character limit (DNS label rules from RFC 1123). The helpers enforce this automatically.</p><h5><strong>Packaging a TypeScript API: the full chart</strong></h5><p>Let&#8217;s build a complete chart for our task API from the series. We need a deployment, a service, an ingress, a configmap, and an HPA. We already saw the ingress above, so let&#8217;s cover the rest.</p><p><strong>Deployment template</strong> (<code>templates/deployment.yaml</code>):</p><pre><code>apiVersion: apps/v1
kind: Deployment
metadata:
  name: {{ include "task-api.fullname" . }}
  labels:
    {{- include "task-api.labels" . | nindent 4 }}
spec:
  {{- if not .Values.autoscaling.enabled }}
  replicas: {{ .Values.replicaCount }}
  {{- end }}
  selector:
    matchLabels:
      {{- include "task-api.selectorLabels" . | nindent 6 }}
  template:
    metadata:
      annotations:
        checksum/config: {{ include (print $.Template.BasePath "/configmap.yaml") . | sha256sum }}
      labels:
        {{- include "task-api.selectorLabels" . | nindent 8 }}
    spec:
      serviceAccountName: {{ include "task-api.serviceAccountName" . }}
      containers:
        - name: {{ .Chart.Name }}
          image: "{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}"
          imagePullPolicy: {{ .Values.image.pullPolicy }}
          ports:
            - name: http
              containerPort: {{ .Values.service.port }}
              protocol: TCP
          envFrom:
            - configMapRef:
                name: {{ include "task-api.fullname" . }}-config
          livenessProbe:
            httpGet:
              path: /health
              port: http
            initialDelaySeconds: 10
            periodSeconds: 15
          readinessProbe:
            httpGet:
              path: /health
              port: http
            initialDelaySeconds: 5
            periodSeconds: 10
          resources:
            {{- toYaml .Values.resources | nindent 12 }}</code></pre><p>Notice the <code>checksum/config</code> annotation on the pod template. This is a common Helm pattern. When you change a ConfigMap, Kubernetes does not automatically restart the pods that use it. By hashing the ConfigMap content into an annotation, any change to the ConfigMap produces a different hash, which triggers a rolling update. Clever and simple.</p><p>Also notice that when autoscaling is enabled, we skip the <code>replicas</code> field. The HPA manages the replica count in that case, and setting it in the deployment would conflict.</p><p><strong>ConfigMap template</strong> (<code>templates/configmap.yaml</code>):</p><pre><code>apiVersion: v1
kind: ConfigMap
metadata:
  name: {{ include "task-api.fullname" . }}-config
  labels:
    {{- include "task-api.labels" . | nindent 4 }}
data:
  {{- range $key, $value := .Values.env }}
  {{ $key }}: {{ $value | quote }}
  {{- end }}</code></pre><p>This loops over every key-value pair in <code>.Values.env</code> and creates a ConfigMap entry. You add new environment variables just by adding them to values.yaml, no template changes needed.</p><p><strong>Service template</strong> (<code>templates/service.yaml</code>):</p><pre><code>apiVersion: v1
kind: Service
metadata:
  name: {{ include "task-api.fullname" . }}
  labels:
    {{- include "task-api.labels" . | nindent 4 }}
spec:
  type: {{ .Values.service.type }}
  ports:
    - port: {{ .Values.service.port }}
      targetPort: http
      protocol: TCP
      name: http
  selector:
    {{- include "task-api.selectorLabels" . | nindent 4 }}</code></pre><p><strong>HPA template</strong> (<code>templates/hpa.yaml</code>):</p><pre><code>{{- if .Values.autoscaling.enabled }}
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: {{ include "task-api.fullname" . }}
  labels:
    {{- include "task-api.labels" . | nindent 4 }}
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: {{ include "task-api.fullname" . }}
  minReplicas: {{ .Values.autoscaling.minReplicas }}
  maxReplicas: {{ .Values.autoscaling.maxReplicas }}
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: {{ .Values.autoscaling.targetCPUUtilizationPercentage }}
{{- end }}</code></pre><p>The entire HPA template is wrapped in an <code>if</code> block. When autoscaling is disabled (the default), this file produces no output at all.</p><h5><strong>Overriding values</strong></h5><p>There are two main ways to override the defaults in values.yaml when you install or upgrade a release.</p><p><strong>Using <code>--set</code> for individual values:</strong></p><pre><code>helm install my-api ./task-api \
  --set image.tag=v1.2.3 \
  --set replicaCount=3 \
  --set ingress.enabled=true</code></pre><p><strong>Using <code>-f</code> (or <code>--values</code>) with a file:</strong></p><pre><code>helm install my-api ./task-api -f production-values.yaml</code></pre><p>Where <code>production-values.yaml</code> might look like:</p><pre><code>replicaCount: 3

image:
  tag: v1.2.3

ingress:
  enabled: true
  className: traefik
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt-prod
  hosts:
    - host: api.example.com
      paths:
        - path: /
          pathType: Prefix
  tls:
    - secretName: api-tls
      hosts:
        - api.example.com

resources:
  limits:
    cpu: 500m
    memory: 512Mi
  requests:
    cpu: 250m
    memory: 256Mi

autoscaling:
  enabled: true
  minReplicas: 3
  maxReplicas: 20
  targetCPUUtilizationPercentage: 70

env:
  NODE_ENV: production
  PORT: "3000"
  LOG_LEVEL: info</code></pre><p>The file approach is better for anything beyond a couple of values. You can version control your environment-specific files (<code>dev-values.yaml</code>, <code>staging-values.yaml</code>, <code>production-values.yaml</code>) and get all the benefits of Git history for configuration changes.</p><p>You can also combine both approaches. Values from <code>-f</code> files are applied first, then <code>--set</code> overrides on top. This is useful when you want a base file plus a one-off override:</p><pre><code>helm install my-api ./task-api \
  -f production-values.yaml \
  --set image.tag=v1.2.4</code></pre><h5><strong>Installing and managing releases</strong></h5><p>Here is the full lifecycle of a Helm release.</p><p><strong>Install a release:</strong></p><pre><code># Install from a local chart directory
helm install my-api ./task-api -n task-api --create-namespace

# Install from a repository
helm install my-api my-repo/task-api -n task-api --create-namespace

# Install and wait for all pods to be ready
helm install my-api ./task-api -n task-api --create-namespace --wait --timeout 5m</code></pre><p>The <code>--wait</code> flag tells Helm to wait until all resources are in a ready state before marking the release as successful. Combined with <code>--timeout</code>, this gives you a clear success or failure signal. Without <code>--wait</code>, Helm marks the release as deployed as soon as the manifests are submitted to the API server, regardless of whether the pods actually start.</p><p><strong>Check release status:</strong></p><pre><code># List all releases in a namespace
helm list -n task-api

# Get detailed status of a release
helm status my-api -n task-api

# See the values that were used for the current release
helm get values my-api -n task-api

# See all values (including defaults)
helm get values my-api -n task-api --all

# See the rendered manifests
helm get manifest my-api -n task-api</code></pre><p><strong>Upgrade a release:</strong></p><pre><code># Upgrade with a new image tag
helm upgrade my-api ./task-api -n task-api --set image.tag=v1.3.0

# Upgrade with a values file
helm upgrade my-api ./task-api -n task-api -f production-values.yaml

# Install or upgrade (idempotent, great for CI/CD)
helm upgrade --install my-api ./task-api -n task-api -f production-values.yaml</code></pre><p>The <code>upgrade --install</code> pattern is the most common in CI/CD pipelines. It installs the release if it does not exist, or upgrades it if it does. This makes your pipeline idempotent, you can run it multiple times without errors.</p><p><strong>View release history:</strong></p><pre><code>helm history my-api -n task-api</code></pre><pre><code>REVISION  UPDATED                   STATUS      CHART           APP VERSION  DESCRIPTION
1         2026-05-24 10:00:00       superseded  task-api-0.1.0  1.0.0        Install complete
2         2026-05-24 14:30:00       superseded  task-api-0.1.0  1.1.0        Upgrade complete
3         2026-05-24 15:00:00       deployed    task-api-0.2.0  1.2.0        Upgrade complete</code></pre><p><strong>Rollback to a previous revision:</strong></p><pre><code># Rollback to the previous revision
helm rollback my-api -n task-api

# Rollback to a specific revision
helm rollback my-api 1 -n task-api</code></pre><p>Rollback is one of the strongest arguments for Helm. If a deployment goes wrong, you can revert to any previous state in seconds. No need to figure out which YAML files to apply or which image tag was running before. Helm tracks all of that for you.</p><p><strong>Uninstall a release:</strong></p><pre><code># Remove the release and all its resources
helm uninstall my-api -n task-api

# Keep the release history (useful for auditing)
helm uninstall my-api -n task-api --keep-history</code></pre><h5><strong>OCI registries: the modern approach</strong></h5><p>Traditional Helm repositories are HTTP servers that host an <code>index.yaml</code> file listing all available charts. They work, but they require maintaining a separate piece of infrastructure.</p><p>The modern approach is to store Helm charts in OCI-compliant container registries, the same registries you already use for Docker images. GitHub Container Registry (GHCR), Docker Hub, ECR, GCR, and Azure Container Registry all support Helm OCI charts.</p><p>Here is how to package and push a chart to GHCR:</p><pre><code># Log in to GHCR
echo $GITHUB_TOKEN | helm registry login ghcr.io --username your-username --password-stdin

# Package the chart
helm package ./task-api

# This creates task-api-0.1.0.tgz in the current directory

# Push to GHCR
helm push task-api-0.1.0.tgz oci://ghcr.io/your-org/charts

# Pull from GHCR
helm pull oci://ghcr.io/your-org/charts/task-api --version 0.1.0

# Install directly from GHCR
helm install my-api oci://ghcr.io/your-org/charts/task-api --version 0.1.0 -n task-api</code></pre><p>The OCI approach has several advantages:</p><blockquote><ul><li><p><strong>No index.yaml</strong>: No need to rebuild and host a chart index. The registry handles discovery.</p></li><li><p><strong>Same infrastructure</strong>: If you already use GHCR for Docker images, you do not need to set up anything else.</p></li><li><p><strong>Access control</strong>: Registry permissions apply to charts the same way they apply to images.</p></li><li><p><strong>Immutable tags</strong>: Once you push a version, it cannot be overwritten (depending on registry settings). This guarantees reproducibility.</p></li></ul></blockquote><p>In a CI/CD pipeline, you would build and push the chart alongside the Docker image:</p><pre><code># .github/workflows/release.yaml (relevant excerpt)
- name: Push Helm chart to GHCR
  run: |
    echo "${{ secrets.GITHUB_TOKEN }}" | helm registry login ghcr.io \
      --username ${{ github.actor }} --password-stdin
    helm package ./charts/task-api --version ${{ github.ref_name }}
    helm push task-api-${{ github.ref_name }}.tgz oci://ghcr.io/${{ github.repository_owner }}/charts</code></pre><h5><strong>Chart testing</strong></h5><p>Before you install a chart in a real cluster, you should validate it. Helm provides several tools for this.</p><p><strong>Linting:</strong></p><pre><code># Check for issues in chart structure and templates
helm lint ./task-api

# Lint with specific values
helm lint ./task-api -f production-values.yaml</code></pre><p><code>helm lint</code> catches common mistakes: missing required fields in Chart.yaml, template syntax errors, indentation problems, and deprecated API versions. Run it in CI on every pull request.</p><p><strong>Template rendering:</strong></p><pre><code># Render templates without installing
helm template my-api ./task-api

# Render with specific values and save to a file for review
helm template my-api ./task-api -f production-values.yaml &gt; rendered.yaml

# Render and validate against the cluster's API
helm template my-api ./task-api --validate</code></pre><p><code>helm template</code> renders the templates locally and prints the resulting YAML. This is incredibly useful for debugging. If something looks wrong in the output, the problem is in your templates or values, not in Kubernetes. The <code>--validate</code> flag adds API server validation, which catches issues like using a removed API version.</p><p><strong>Release testing:</strong></p><pre><code># Run the chart's test pods
helm test my-api -n task-api</code></pre><p>Helm supports test hooks. These are pods defined in <code>templates/tests/</code> that run when you execute <code>helm test</code>. A typical test verifies that the deployed application is reachable:</p><pre><code># templates/tests/test-connection.yaml
apiVersion: v1
kind: Pod
metadata:
  name: "{{ include "task-api.fullname" . }}-test-connection"
  labels:
    {{- include "task-api.labels" . | nindent 4 }}
  annotations:
    "helm.sh/hook": test
spec:
  containers:
    - name: wget
      image: busybox
      command: ['wget']
      args: ['{{ include "task-api.fullname" . }}:{{ .Values.service.port }}/health']
  restartPolicy: Never</code></pre><p>This pod runs <code>wget</code> against the service&#8217;s health endpoint. If it succeeds, the test passes. If it fails, you know something is wrong with the deployment.</p><h5><strong>Managing multiple charts: Helmfile and ArgoCD</strong></h5><p>Once you have more than a handful of charts, you need a way to manage them together. Two tools stand out.</p><p><strong>Helmfile</strong> is a declarative spec for deploying multiple Helm charts. Instead of running <code>helm install</code> and <code>helm upgrade</code> commands manually, you define everything in a <code>helmfile.yaml</code>:</p><pre><code># helmfile.yaml
repositories:
  - name: bitnami
    url: https://charts.bitnami.com/bitnami

releases:
  - name: task-api
    namespace: task-api
    chart: ./charts/task-api
    values:
      - environments/{{ .Environment.Name }}/task-api.yaml

  - name: redis
    namespace: task-api
    chart: bitnami/redis
    version: 18.6.1
    values:
      - environments/{{ .Environment.Name }}/redis.yaml</code></pre><p>Then deploy everything with:</p><pre><code>helmfile -e production apply</code></pre><p><strong>ArgoCD</strong> takes a different approach. Instead of running commands, you define your desired state in Git and ArgoCD continuously reconciles the cluster to match. ArgoCD has native Helm support, so you point it at a Git repository containing your chart and values, and it handles the rest:</p><pre><code># ArgoCD Application manifest
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: task-api
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/your-org/your-repo
    targetRevision: main
    path: charts/task-api
    helm:
      valueFiles:
        - ../../environments/production/task-api.yaml
  destination:
    server: https://kubernetes.default.svc
    namespace: task-api
  syncPolicy:
    automated:
      selfHeal: true
      prune: true
    syncOptions:
      - CreateNamespace=true</code></pre><p>ArgoCD is the GitOps approach. Every change goes through a pull request, gets reviewed, merged to main, and ArgoCD applies it automatically. No one runs <code>helm install</code> or <code>kubectl apply</code> manually. This is how most production teams operate today. If you want to dig deeper into ArgoCD, check out <a href="https://segfault.pw/blog/sre-gitops-with-argocd">GitOps with ArgoCD</a> from the SRE series.</p><h5><strong>Common patterns and tips</strong></h5><p>Here are some patterns you will encounter often when working with Helm.</p><p><strong>Use <code>helm upgrade --install</code> in CI/CD.</strong> This makes deployments idempotent. Whether the release exists or not, the command does the right thing.</p><p><strong>Always set resource requests and limits.</strong> The default values.yaml should include reasonable resource values. Without them, a single pod can consume all cluster resources.</p><p><strong>Use the checksum annotation pattern.</strong> As we saw earlier, hashing ConfigMap content into a pod annotation triggers rolling updates when configuration changes. This saves you from the &#8220;I changed the ConfigMap but nothing happened&#8221; surprise.</p><p><strong>Pin your chart versions.</strong> When installing from a repository, always specify <code>--version</code>. Without it, Helm installs the latest version, which might introduce breaking changes.</p><p><strong>Keep secrets out of values.yaml.</strong> Never put passwords, API keys, or tokens in values files that get committed to Git. Use Kubernetes Secrets managed by an external tool like External Secrets Operator or Sealed Secrets.</p><p><strong>Use <code>helm diff</code> for safe upgrades.</strong> The <code>helm-diff</code> plugin shows you exactly what will change before you upgrade:</p><pre><code># Install the diff plugin
helm plugin install https://github.com/databus23/helm-diff

# Preview changes before upgrading
helm diff upgrade my-api ./task-api -f production-values.yaml -n task-api</code></pre><p>This is especially valuable in production where you want to review changes before applying them.</p><h5><strong>Closing notes</strong></h5><p>Helm takes the pain out of managing Kubernetes applications. Instead of juggling raw YAML files across environments, you define your application once as a chart, parameterize the things that change, and let Helm handle the rendering, versioning, and lifecycle management.</p><p>In this article we covered what Helm is and why it exists, created a chart from scratch with all the templates a real application needs, explored Go template syntax including conditionals, loops, and built-in objects, built reusable helpers, managed releases with install, upgrade, rollback, and history, pushed charts to OCI registries for modern distribution, and validated everything with lint, template, and test.</p><p>The key takeaway is that Helm is not just about templating YAML. It is about giving your Kubernetes deployments a proper lifecycle: versioned releases, configuration management, rollback capability, and a shared language for your team to talk about what is running where.</p><p>In the next article we will look at monitoring and observability for our Kubernetes workloads, because deploying an application is only half the job. You also need to know if it is healthy and performing well.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Kubernetes Fundamentals]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-kubernetes-fundamentals</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-kubernetes-fundamentals</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!H51I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!H51I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!H51I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!H51I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!H51I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!H51I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!H51I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043136?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!H51I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!H51I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!H51I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!H51I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1de1c3ef-5525-43da-bf42-53619a508936_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article eleven of the DevOps from Zero to Hero series. In article eight we deployed our TypeScript API to AWS ECS with Fargate. ECS is a solid container orchestrator, but it is AWS-specific. If you want something that runs on any cloud provider, on bare metal, or even on your laptop, Kubernetes is the answer.</p><p>Kubernetes (often shortened to K8s) is the industry standard for container orchestration. It is what most teams end up using when they need to run containers at scale. It is also one of those technologies that looks intimidating from the outside but makes a lot of sense once you understand the core concepts.</p><p>In this article we will cover what Kubernetes is and why it exists, walk through the architecture, learn about every core object you will use daily, set up a local cluster with kind, and deploy a real workload step by step. By the end you will be comfortable reading Kubernetes manifests, running kubectl commands, and understanding what is happening inside a cluster.</p><p>Let&#8217;s get into it.</p><h5><strong>What is Kubernetes and why does it exist?</strong></h5><p>Imagine you have ten containers that need to run across five servers. Some containers need to talk to each other. Some need more CPU than others. If one crashes, you want it restarted automatically. If traffic spikes, you want to spin up more copies. And you want all of this to happen without you waking up at 3 AM.</p><p>That is container orchestration, and that is what Kubernetes does. It takes a set of machines, pools their resources together, and lets you declare what you want running. Kubernetes then figures out where to place each container, keeps everything healthy, and handles networking so containers can find each other.</p><p>The key capabilities are:</p><blockquote><ul><li><p><strong>Scheduling</strong>: Kubernetes decides which node (server) each container runs on based on available resources.</p></li><li><p><strong>Scaling</strong>: You tell Kubernetes how many copies of a container you want. It makes it happen. You can also set up auto-scaling based on CPU, memory, or custom metrics.</p></li><li><p><strong>Self-healing</strong>: If a container crashes, Kubernetes restarts it. If a node goes down, Kubernetes reschedules the containers that were running on it to healthy nodes.</p></li><li><p><strong>Service discovery and load balancing</strong>: Kubernetes gives each set of containers a stable network identity and balances traffic across them automatically.</p></li><li><p><strong>Rolling updates and rollbacks</strong>: You can update your application with zero downtime. If something goes wrong, you can roll back to the previous version with a single command.</p></li><li><p><strong>Declarative configuration</strong>: You describe what you want in YAML files, and Kubernetes continuously works to make reality match your description. This is called the &#8220;desired state&#8221; model.</p></li></ul></blockquote><p>Kubernetes was originally designed by Google, based on their internal system called Borg. It was open sourced in 2014 and is now maintained by the Cloud Native Computing Foundation (CNCF). Every major cloud provider offers a managed Kubernetes service: EKS on AWS, GKE on Google Cloud, AKS on Azure.</p><h5><strong>Architecture overview</strong></h5><p>A Kubernetes cluster has two types of components: the control plane (the brain) and the worker nodes (the muscle). Here is how they fit together:</p><pre><code>+-----------------------------------------------------------+
|                     Kubernetes Cluster                     |
|                                                           |
|  +-----------------------------------------------------+  |
|  |                   Control Plane                      |  |
|  |                                                     |  |
|  |  +--------------+  +-------+  +-----------+         |  |
|  |  |  API Server  |  | etcd  |  | Scheduler |         |  |
|  |  +--------------+  +-------+  +-----------+         |  |
|  |  +--------------------+                             |  |
|  |  | Controller Manager |                             |  |
|  |  +--------------------+                             |  |
|  +-----------------------------------------------------+  |
|                                                           |
|  +------------------------+  +------------------------+   |
|  |     Worker Node 1      |  |     Worker Node 2      |   |
|  |                        |  |                        |   |
|  |  +--------+ +-------+ |  |  +--------+ +-------+  |   |
|  |  | kubelet| | proxy | |  |  | kubelet| | proxy |  |   |
|  |  +--------+ +-------+ |  |  +--------+ +-------+  |   |
|  |  +------+ +------+    |  |  +------+ +------+     |   |
|  |  | Pod  | | Pod  |    |  |  | Pod  | | Pod  |     |   |
|  |  +------+ +------+    |  |  +------+ +------+     |   |
|  +------------------------+  +------------------------+   |
+-----------------------------------------------------------+</code></pre><p>Let&#8217;s break down each component:</p><h5><strong>Control plane components</strong></h5><blockquote><ul><li><p><strong>API Server (kube-apiserver)</strong>: The front door to your cluster. Every command you run with kubectl goes through the API server. It validates requests, updates the cluster state, and is the only component that talks directly to etcd. Think of it as the receptionist that handles all incoming requests.</p></li><li><p><strong>etcd</strong>: A distributed key-value store that holds the entire state of your cluster. Every object you create, every configuration, every secret is stored here. If etcd is lost and you have no backup, your cluster state is gone. It is the single source of truth.</p></li><li><p><strong>Scheduler (kube-scheduler)</strong>: When you create a new Pod and it does not have a node assigned yet, the scheduler picks one. It looks at resource requirements, constraints, and available capacity to make the best placement decision.</p></li><li><p><strong>Controller Manager (kube-controller-manager)</strong>: Runs a set of controllers that watch the cluster state and work to make reality match the desired state. For example, the ReplicaSet controller ensures the right number of Pod replicas are running. If you ask for three replicas and only two are running, it creates another one.</p></li></ul></blockquote><h5><strong>Worker node components</strong></h5><blockquote><ul><li><p><strong>kubelet</strong>: An agent that runs on every worker node. It receives Pod specifications from the API server and ensures the containers described in those specs are running and healthy. If a container crashes, kubelet restarts it.</p></li><li><p><strong>kube-proxy</strong>: Manages network rules on each node. It handles the networking magic that lets you reach any Pod from any node using a Service. It sets up iptables rules (or IPVS, depending on configuration) to route traffic correctly.</p></li><li><p><strong>Container runtime</strong>: The software that actually runs containers. Kubernetes supports any runtime that implements the Container Runtime Interface (CRI). The most common ones are containerd and CRI-O. Docker used to be the default, but Kubernetes removed direct Docker support in version 1.24 (containerd, which Docker uses under the hood, is still fully supported).</p></li></ul></blockquote><h5><strong>Core objects: Pods</strong></h5><p>A Pod is the smallest deployable unit in Kubernetes. It is not a container. It is a wrapper around one or more containers that share the same network namespace and storage volumes.</p><p>Most of the time a Pod runs a single container. But sometimes you need a helper container alongside your main one (for logging, proxying, or injecting configuration). Those are called sidecar containers, and they live in the same Pod.</p><p>Containers in the same Pod:</p><blockquote><ul><li><p><strong>Share the same IP address</strong> and can talk to each other via localhost</p></li><li><p><strong>Share storage volumes</strong> mounted into the Pod</p></li><li><p><strong>Are scheduled together</strong> on the same node</p></li><li><p><strong>Start and stop together</strong> as a unit</p></li></ul></blockquote><p>Here is a simple Pod definition:</p><pre><code>apiVersion: v1
kind: Pod
metadata:
  name: my-nginx
  labels:
    app: nginx
spec:
  containers:
    - name: nginx
      image: nginx:1.27
      ports:
        - containerPort: 80</code></pre><p>You almost never create Pods directly in production. Instead, you use a Deployment (which we will cover next) that manages Pods for you. If you create a Pod directly and it crashes, nothing will restart it. A Deployment ensures crashed Pods are replaced automatically.</p><h5><strong>Core objects: Deployments</strong></h5><p>A Deployment is the most common way to run workloads in Kubernetes. It wraps a Pod template and adds powerful management features on top.</p><p>When you create a Deployment, you tell Kubernetes: &#8220;I want three replicas of this container, always running, and here is how to update them.&#8221; Kubernetes then creates a ReplicaSet behind the scenes, and the ReplicaSet creates the Pods. The chain looks like this:</p><pre><code>Deployment
  &#9492;&#9472;&#9472; ReplicaSet
        &#9500;&#9472;&#9472; Pod 1
        &#9500;&#9472;&#9472; Pod 2
        &#9492;&#9472;&#9472; Pod 3</code></pre><p>Here is a Deployment manifest:</p><pre><code>apiVersion: apps/v1
kind: Deployment
metadata:
  name: nginx-deployment
  labels:
    app: nginx
spec:
  replicas: 3
  selector:
    matchLabels:
      app: nginx
  template:
    metadata:
      labels:
        app: nginx
    spec:
      containers:
        - name: nginx
          image: nginx:1.27
          ports:
            - containerPort: 80
          resources:
            requests:
              memory: "64Mi"
              cpu: "100m"
            limits:
              memory: "128Mi"
              cpu: "250m"</code></pre><p>Key features of Deployments:</p><blockquote><ul><li><p><strong>Desired state</strong>: You declare how many replicas you want. If a Pod dies, the Deployment creates a new one. If you have too many, it terminates the extras.</p></li><li><p><strong>Rolling updates</strong>: When you change the container image, the Deployment gradually replaces old Pods with new ones, ensuring zero downtime. By default it takes down at most 25% of Pods at a time while bringing up new ones.</p></li><li><p><strong>Rollback</strong>: Every change to a Deployment creates a new revision. If a new version is broken, you can roll back to any previous revision with <code>kubectl rollout undo</code>.</p></li><li><p><strong>Scaling</strong>: Change the replica count and Kubernetes handles the rest. Scale up or down at any time.</p></li></ul></blockquote><h5><strong>Core objects: Services</strong></h5><p>Pods are ephemeral. They get created, destroyed, and moved around constantly. Each time a Pod is recreated, it gets a new IP address. So how do other Pods find and talk to your application?</p><p>That is what Services solve. A Service provides a stable network endpoint (a fixed IP and DNS name) that routes traffic to a set of Pods. Even as Pods come and go, the Service keeps pointing to the healthy ones.</p><p>There are three main types:</p><blockquote><ul><li><p><strong>ClusterIP (default)</strong>: Creates an internal IP address that is only reachable from within the cluster. This is what you use for service-to-service communication. For example, your API talking to your database.</p></li><li><p><strong>NodePort</strong>: Exposes the service on a static port on every node in the cluster. You can reach it from outside by hitting any node&#8217;s IP at that port. Useful for development, but not ideal for production.</p></li><li><p><strong>LoadBalancer</strong>: Provisions an external load balancer (on cloud providers). This is the standard way to expose a service to the internet in production. On AWS it creates an ELB, on GCP a Cloud Load Balancer, and so on.</p></li></ul></blockquote><p>Here is a Service that exposes our nginx Deployment:</p><pre><code>apiVersion: v1
kind: Service
metadata:
  name: nginx-service
spec:
  type: ClusterIP
  selector:
    app: nginx
  ports:
    - protocol: TCP
      port: 80
      targetPort: 80</code></pre><p>The <code>selector</code> field is what connects a Service to its Pods. The Service looks for all Pods with the label <code>app: nginx</code> and routes traffic to them. This is the label-selector mechanism and it is fundamental to how Kubernetes connects objects together.</p><h5><strong>Core objects: ConfigMaps and Secrets</strong></h5><p>Applications need configuration: database URLs, feature flags, API keys. Hardcoding these values into your container image is a bad idea because you would need to rebuild the image for every environment.</p><p>Kubernetes solves this with ConfigMaps and Secrets:</p><blockquote><ul><li><p><strong>ConfigMap</strong>: Stores non-sensitive configuration as key-value pairs. Things like environment names, log levels, and feature flags.</p></li><li><p><strong>Secret</strong>: Stores sensitive data like passwords, tokens, and certificates. Secrets are base64-encoded (not encrypted by default, but you can enable encryption at rest). In production, use a secrets manager like HashiCorp Vault or AWS Secrets Manager and sync secrets into Kubernetes with an operator.</p></li></ul></blockquote><p>Here is a ConfigMap:</p><pre><code>apiVersion: v1
kind: ConfigMap
metadata:
  name: app-config
data:
  LOG_LEVEL: "info"
  APP_ENV: "production"
  MAX_CONNECTIONS: "100"</code></pre><p>And a Secret:</p><pre><code>apiVersion: v1
kind: Secret
metadata:
  name: app-secrets
type: Opaque
data:
  DATABASE_URL: cG9zdGdyZXM6Ly91c2VyOnBhc3NAaG9zdDo1NDMyL2Ri
  API_KEY: c3VwZXItc2VjcmV0LWtleQ==</code></pre><p>You can inject these into Pods as environment variables or mount them as files. Here is how to use both in a Deployment:</p><pre><code>spec:
  containers:
    - name: app
      image: my-app:1.0
      envFrom:
        - configMapRef:
            name: app-config
        - secretRef:
            name: app-secrets</code></pre><h5><strong>Core objects: Namespaces</strong></h5><p>Namespaces provide logical isolation within a cluster. They are like folders for your Kubernetes objects. Different teams, environments, or applications can each have their own namespace.</p><p>Every cluster starts with a few default namespaces:</p><blockquote><ul><li><p><strong>default</strong>: Where objects go if you do not specify a namespace.</p></li><li><p><strong>kube-system</strong>: Where Kubernetes system components run (API server, scheduler, CoreDNS, etc.).</p></li><li><p><strong>kube-public</strong>: Readable by all users, used for cluster-wide public information.</p></li><li><p><strong>kube-node-lease</strong>: Holds lease objects for node heartbeats.</p></li></ul></blockquote><p>Creating a namespace is simple:</p><pre><code>kubectl create namespace staging</code></pre><p>Or with YAML:</p><pre><code>apiVersion: v1
kind: Namespace
metadata:
  name: staging</code></pre><p>Then deploy resources into that namespace:</p><pre><code>kubectl apply -f deployment.yaml -n staging</code></pre><p>Namespaces are also the boundary for resource quotas and network policies. You can limit how much CPU and memory a namespace can consume, and you can control which namespaces can talk to each other.</p><h5><strong>Labels and selectors</strong></h5><p>Labels are key-value pairs attached to any Kubernetes object. They are the glue that connects different objects together.</p><pre><code>metadata:
  labels:
    app: nginx
    environment: production
    team: platform
    version: "1.27"</code></pre><p>Selectors filter objects based on their labels. This is how a Service finds its Pods, how a Deployment knows which Pods it owns, and how you can query specific objects with kubectl:</p><pre><code># Get all pods with a specific label
kubectl get pods -l app=nginx

# Get pods matching multiple labels
kubectl get pods -l app=nginx,environment=production

# Get pods where a label exists (any value)
kubectl get pods -l team

# Get pods where a label does NOT exist
kubectl get pods -l '!team'</code></pre><p>Labels and selectors are not just a nice-to-have. They are how Kubernetes works internally. If your Service selector does not match your Pod labels, traffic will not flow. If your Deployment selector does not match the Pod template labels, the Deployment will reject the configuration.</p><h5><strong>Resource requests and limits</strong></h5><p>Every container should declare how much CPU and memory it needs. Without this, Kubernetes has no idea how to schedule Pods efficiently and you risk overloading nodes.</p><p>There are two settings:</p><blockquote><ul><li><p><strong>Requests</strong>: The minimum amount of resources guaranteed to the container. The scheduler uses requests to decide which node has enough room for the Pod. If you request 256Mi of memory, Kubernetes will place the Pod on a node with at least that much available.</p></li><li><p><strong>Limits</strong>: The maximum amount of resources a container can use. If a container exceeds its memory limit, Kubernetes kills it (OOMKilled). If it exceeds its CPU limit, it gets throttled (slowed down but not killed).</p></li></ul></blockquote><pre><code>resources:
  requests:
    memory: "128Mi"
    cpu: "100m"
  limits:
    memory: "256Mi"
    cpu: "500m"</code></pre><p>CPU is measured in millicores. <code>100m</code> means 0.1 CPU cores. <code>1000m</code> (or just <code>1</code>) means one full core. Memory uses standard units: <code>Mi</code> (mebibytes), <code>Gi</code> (gibibytes).</p><p>A few rules of thumb:</p><blockquote><ul><li><p><strong>Always set requests</strong>. Without them, the scheduler is guessing.</p></li><li><p><strong>Set memory limits</strong> to prevent runaway containers from crashing the node.</p></li><li><p><strong>Be careful with CPU limits</strong>. Aggressive CPU limits cause throttling even when the node has spare CPU. Some teams set CPU requests but skip CPU limits to avoid unnecessary throttling.</p></li><li><p><strong>Monitor actual usage</strong> and adjust requests/limits based on real data, not guesses.</p></li></ul></blockquote><h5><strong>Setting up a local cluster with kind</strong></h5><p>kind (Kubernetes in Docker) is the fastest way to get a local Kubernetes cluster running. It creates a cluster by running Kubernetes nodes as Docker containers. You need Docker installed and that is it.</p><p>Install kind:</p><pre><code># On Linux
curl -Lo ./kind https://kind.sigs.k8s.io/dl/v0.27.0/kind-linux-amd64
chmod +x ./kind
sudo mv ./kind /usr/local/bin/kind

# On macOS (Homebrew)
brew install kind</code></pre><p>Create a cluster:</p><pre><code>kind create cluster --name my-cluster</code></pre><p>That is it. kind creates a single-node cluster and configures kubectl to use it. Verify it is running:</p><pre><code>kubectl cluster-info --context kind-my-cluster
kubectl get nodes</code></pre><p>You should see output like:</p><pre><code>NAME                       STATUS   ROLES           AGE   VERSION
my-cluster-control-plane   Ready    control-plane   45s   v1.32.2</code></pre><p>When you are done, delete the cluster:</p><pre><code>kind delete cluster --name my-cluster</code></pre><h5><strong>kubectl basics</strong></h5><p>kubectl is the command-line tool for interacting with Kubernetes. Here are the commands you will use every day:</p><pre><code># Get resources
kubectl get pods                     # List all pods in current namespace
kubectl get pods -A                  # List pods in ALL namespaces
kubectl get deployments              # List deployments
kubectl get services                 # List services
kubectl get all                      # List common resource types

# Detailed information about a resource
kubectl describe pod my-nginx        # Show events, conditions, containers
kubectl describe deployment nginx-deployment

# View logs
kubectl logs my-nginx                # Logs from a pod
kubectl logs my-nginx -f             # Stream logs (follow)
kubectl logs my-nginx --previous     # Logs from the previous container (after crash)

# Execute commands inside a container
kubectl exec -it my-nginx -- /bin/bash   # Interactive shell
kubectl exec my-nginx -- cat /etc/nginx/nginx.conf  # Run a single command

# Apply and delete resources from files
kubectl apply -f deployment.yaml     # Create or update resources from a file
kubectl apply -f ./manifests/        # Apply all files in a directory
kubectl delete -f deployment.yaml    # Delete resources defined in a file
kubectl delete pod my-nginx          # Delete a specific pod</code></pre><p>A few tips that will save you time:</p><blockquote><ul><li><p><strong>Use <code>-o wide</code></strong> to see extra columns like node name and IP: <code>kubectl get pods -o wide</code></p></li><li><p><strong>Use <code>-o yaml</code></strong> to see the full object definition: <code>kubectl get pod my-nginx -o yaml</code></p></li><li><p><strong>Set a default namespace</strong> so you do not have to type <code>-n</code> every time: <code>kubectl config set-context --current --namespace=staging</code></p></li><li><p><strong>Use aliases</strong>. Most Kubernetes users alias <code>kubectl</code> to <code>k</code>: <code>alias k=kubectl</code></p></li></ul></blockquote><h5><strong>Practical walkthrough: deploy, expose, scale, update</strong></h5><p>Let&#8217;s put everything together with a hands-on exercise. Make sure you have a kind cluster running.</p><p><strong>Step 1: Create a Deployment</strong></p><p>Create a file called <code>nginx-deployment.yaml</code>:</p><pre><code>apiVersion: apps/v1
kind: Deployment
metadata:
  name: nginx-deployment
  labels:
    app: nginx
spec:
  replicas: 2
  selector:
    matchLabels:
      app: nginx
  template:
    metadata:
      labels:
        app: nginx
    spec:
      containers:
        - name: nginx
          image: nginx:1.27
          ports:
            - containerPort: 80
          resources:
            requests:
              memory: "64Mi"
              cpu: "50m"
            limits:
              memory: "128Mi"
              cpu: "100m"</code></pre><p>Apply it:</p><pre><code>kubectl apply -f nginx-deployment.yaml</code></pre><p>Check the results:</p><pre><code>kubectl get deployments
kubectl get pods</code></pre><p>You should see two Pods running:</p><pre><code>NAME                                READY   STATUS    RESTARTS   AGE
nginx-deployment-5d8f4d7b9c-abc12   1/1     Running   0          15s
nginx-deployment-5d8f4d7b9c-def34   1/1     Running   0          15s</code></pre><p><strong>Step 2: Expose it with a Service</strong></p><p>Create a file called <code>nginx-service.yaml</code>:</p><pre><code>apiVersion: v1
kind: Service
metadata:
  name: nginx-service
spec:
  type: NodePort
  selector:
    app: nginx
  ports:
    - protocol: TCP
      port: 80
      targetPort: 80
      nodePort: 30080</code></pre><p>Apply it:</p><pre><code>kubectl apply -f nginx-service.yaml</code></pre><p>Verify the Service:</p><pre><code>kubectl get services</code></pre><pre><code>NAME            TYPE       CLUSTER-IP     EXTERNAL-IP   PORT(S)        AGE
nginx-service   NodePort   10.96.45.123   &lt;none&gt;        80:30080/TCP   5s
kubernetes      ClusterIP  10.96.0.1      &lt;none&gt;        443/TCP        10m</code></pre><p>Test that it works. Since we are using kind, we can port-forward to access the service locally:</p><pre><code>kubectl port-forward service/nginx-service 8080:80</code></pre><p>Now open another terminal and hit it:</p><pre><code>curl http://localhost:8080</code></pre><p>You should see the default nginx welcome page HTML.</p><p><strong>Step 3: Scale the Deployment</strong></p><p>Let&#8217;s go from two replicas to five:</p><pre><code>kubectl scale deployment nginx-deployment --replicas=5</code></pre><p>Watch the Pods come up:</p><pre><code>kubectl get pods -w</code></pre><p>Within seconds you will have five Pods running. Scale back down:</p><pre><code>kubectl scale deployment nginx-deployment --replicas=2</code></pre><p>Kubernetes will terminate three Pods gracefully.</p><p><strong>Step 4: Do a rolling update</strong></p><p>Let&#8217;s update from nginx 1.27 to 1.28. You can edit the YAML file and re-apply, or do it inline:</p><pre><code>kubectl set image deployment/nginx-deployment nginx=nginx:1.28</code></pre><p>Watch the rolling update happen:</p><pre><code>kubectl rollout status deployment/nginx-deployment</code></pre><pre><code>Waiting for deployment "nginx-deployment" rollout to finish: 1 out of 2 new replicas have been updated...
Waiting for deployment "nginx-deployment" rollout to finish: 1 old replicas are pending termination...
deployment "nginx-deployment" successfully rolled out</code></pre><p>Kubernetes created new Pods with nginx 1.28 and terminated the old ones, one at a time, with zero downtime.</p><p>Check the rollout history:</p><pre><code>kubectl rollout history deployment/nginx-deployment</code></pre><p>If something goes wrong, roll back:</p><pre><code>kubectl rollout undo deployment/nginx-deployment</code></pre><p>This reverts to the previous revision immediately.</p><p><strong>Step 5: Inspect and debug</strong></p><p>Get detailed information about a Pod:</p><pre><code>kubectl describe pod nginx-deployment-&lt;tab-complete-the-name&gt;</code></pre><p>Check the container logs:</p><pre><code>kubectl logs deployment/nginx-deployment</code></pre><p>Open a shell inside a running container:</p><pre><code>kubectl exec -it deployment/nginx-deployment -- /bin/bash</code></pre><p>Inside the container you can inspect files, test connectivity, and debug issues directly.</p><p><strong>Step 6: Clean up</strong></p><pre><code>kubectl delete -f nginx-service.yaml
kubectl delete -f nginx-deployment.yaml</code></pre><p>Or delete the entire kind cluster:</p><pre><code>kind delete cluster --name my-cluster</code></pre><h5><strong>Closing notes</strong></h5><p>Kubernetes has a reputation for being complex, and it is true that the ecosystem is massive. But the core concepts are straightforward. You have Pods that run containers, Deployments that manage Pods, Services that route traffic, ConfigMaps and Secrets for configuration, and Namespaces for isolation. Everything connects through labels and selectors.</p><p>The key insight is that Kubernetes is a declarative system. You tell it what you want, and it continuously works to make that happen. You do not tell it &#8220;start three containers.&#8221; You tell it &#8220;I want three replicas&#8221; and it figures out how to get there, whether that means creating new Pods, restarting crashed ones, or rescheduling them to different nodes.</p><p>We covered a lot of ground in this article. Set up a kind cluster and play around. Break things on purpose. Delete a Pod and watch the Deployment recreate it. Change resource limits and see what happens. The best way to learn Kubernetes is by using it.</p><p>In the next articles we will build on this foundation: deploying real applications to Kubernetes, setting up networking with Ingress controllers, and managing everything with Helm charts.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: DNS, TLS, and Making Your App Reachable]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-dns-tls-and-networking</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-dns-tls-and-networking</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Mon, 18 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kJHz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kJHz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kJHz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!kJHz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!kJHz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!kJHz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kJHz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043137?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kJHz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!kJHz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!kJHz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!kJHz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce342d35-2c66-47b8-bbe5-23eba26761e4_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article ten of the DevOps from Zero to Hero series. In article eight we deployed our TypeScript API to ECS with Fargate and put an Application Load Balancer in front of it. That setup works, but right now the only way to reach the API is through an ugly auto-generated ALB hostname like <code>task-api-alb-123456789.us-east-1.elb.amazonaws.com</code>. No one wants to type that into a browser, and no one should be sending real traffic over plain HTTP.</p><p>In this article we are going to fix both of those problems. We will register a domain name, set up DNS so that a clean URL like <code>api.yourdomain.com</code> points to our load balancer, and configure TLS so all traffic is encrypted with HTTPS. By the end, your application will be reachable at a real domain with a valid certificate, exactly like a production service should be.</p><p>Let&#8217;s get into it.</p><h5><strong>DNS fundamentals: how your browser finds a server</strong></h5><p>DNS (Domain Name System) is the phonebook of the internet. When you type <code>google.com</code> into your browser, your computer does not magically know which IP address to connect to. It asks the DNS system, and DNS translates that human-readable name into a machine-readable IP address.</p><p>Here is the simplified flow of what happens when you visit <code>api.yourdomain.com</code>:</p><pre><code>1. Browser asks: "What is the IP for api.yourdomain.com?"
2. Your OS checks its local cache. Not found.
3. Your OS asks your ISP's recursive resolver.
4. Resolver asks a root nameserver: "Who handles .com?"
5. Root says: "Ask the .com TLD nameserver."
6. Resolver asks the .com TLD: "Who handles yourdomain.com?"
7. TLD says: "Ask ns-1234.awsdns-56.org (the authoritative nameserver)."
8. Resolver asks the authoritative nameserver: "What is api.yourdomain.com?"
9. Authoritative nameserver responds: "It is an A record pointing to 54.23.45.67."
10. Browser connects to 54.23.45.67.</code></pre><p>This entire process usually takes less than 100 milliseconds. Once resolved, the result is cached at multiple levels so subsequent requests are almost instant.</p><h5><strong>DNS record types you need to know</strong></h5><p>DNS is not just about mapping names to IP addresses. There are several record types, each serving a different purpose:</p><blockquote><ul><li><p><strong>A record</strong>: Maps a domain name to an IPv4 address. Example: <code>api.yourdomain.com -&gt; 54.23.45.67</code>. This is the most basic record type.</p></li><li><p><strong>AAAA record</strong>: Same as A, but for IPv6 addresses. Example: <code>api.yourdomain.com -&gt; 2600:1f18:243:...</code>. As IPv6 adoption grows, you will see more of these.</p></li><li><p><strong>CNAME record</strong>: Maps a domain name to another domain name (an alias). Example: <code>www.yourdomain.com -&gt; yourdomain.com</code>. The resolver follows the chain until it reaches an A record. Important: you cannot use a CNAME at the zone apex (the bare domain like <code>yourdomain.com</code>).</p></li><li><p><strong>NS record</strong>: Specifies the authoritative nameservers for a domain. When you register a domain, the registrar needs to know which nameservers hold the DNS records for that domain.</p></li><li><p><strong>MX record</strong>: Specifies mail servers for a domain. Not relevant for this article, but you will encounter these if you ever set up email.</p></li><li><p><strong>TXT record</strong>: Holds arbitrary text. Used for domain verification (proving you own a domain), SPF records for email, and TLS certificate validation (which we will use later).</p></li></ul></blockquote><h5><strong>TTL: how long DNS answers are cached</strong></h5><p>Every DNS record has a TTL (Time To Live), measured in seconds. It tells resolvers how long to cache the answer before asking again.</p><blockquote><ul><li><p><strong>High TTL (86400 = 24 hours)</strong>: Good for records that rarely change. Reduces DNS queries, faster for users. Bad for quick changes since you have to wait for caches to expire.</p></li><li><p><strong>Low TTL (60 = 1 minute)</strong>: Good when you expect to change records frequently (during migrations, failovers). More DNS queries, slightly higher latency on first request.</p></li><li><p><strong>Common strategy</strong>: Keep TTL high for normal operations. Before a planned migration, lower the TTL a day or two in advance, do the change, verify it works, then raise the TTL back up.</p></li></ul></blockquote><p>A common mistake is forgetting about TTL when doing a migration. If your TTL is 24 hours and you change an A record, some users will still be hitting the old IP for up to 24 hours. Plan ahead.</p><h5><strong>Route53: AWS&#8217;s DNS service</strong></h5><p>Route53 is AWS&#8217;s managed DNS service. It is highly available, globally distributed, and integrates natively with other AWS services. The name comes from the fact that DNS uses port 53.</p><p>The core concept in Route53 is the <strong>hosted zone</strong>. A hosted zone is a container for DNS records that belong to a single domain. When you create a hosted zone for <code>yourdomain.com</code>, Route53 assigns four nameservers to it. You then configure your domain registrar to point at those nameservers.</p><p>There are two types of hosted zones:</p><blockquote><ul><li><p><strong>Public hosted zone</strong>: Resolves queries from the public internet. This is what you need for your user-facing application.</p></li><li><p><strong>Private hosted zone</strong>: Resolves queries only within a VPC. Useful for internal service discovery (e.g., <code>database.internal.yourdomain.com</code>).</p></li></ul></blockquote><h5><strong>Route53 routing policies</strong></h5><p>Route53 supports several routing policies that go beyond simple &#8220;name to IP&#8221; resolution:</p><blockquote><ul><li><p><strong>Simple routing</strong>: One record, one or more values. Route53 returns all values in random order. This is the default and what we will use in this article.</p></li><li><p><strong>Weighted routing</strong>: Distribute traffic across multiple resources by weight. For example, send 90% of traffic to version 1 and 10% to version 2. Great for canary deployments.</p></li><li><p><strong>Failover routing</strong>: Active-passive setup. Route53 sends traffic to the primary resource. If a health check fails, it automatically switches to a secondary resource.</p></li><li><p><strong>Latency-based routing</strong>: Route users to the region with the lowest latency. If you have servers in us-east-1 and eu-west-1, European users automatically get routed to eu-west-1.</p></li><li><p><strong>Geolocation routing</strong>: Route based on the user&#8217;s geographic location. Useful for compliance (keep EU user data in the EU) or serving localized content.</p></li></ul></blockquote><p>For most applications starting out, simple routing is all you need. As you grow and deploy to multiple regions, the other policies become incredibly valuable.</p><h5><strong>Domain registration and nameserver delegation</strong></h5><p>Before you can use Route53 for DNS, you need a domain name. You have two options:</p><blockquote><ul><li><p><strong>Register through Route53</strong>: AWS acts as both your registrar and DNS provider. This is the simplest option because the nameserver delegation happens automatically.</p></li><li><p><strong>Register elsewhere and delegate to Route53</strong>: Buy your domain from a registrar like Namecheap, GoDaddy, or Cloudflare, then update the nameservers to point at the Route53 hosted zone&#8217;s NS records.</p></li></ul></blockquote><p>If you registered your domain elsewhere, the process looks like this:</p><pre><code>1. Create a hosted zone in Route53 for yourdomain.com
2. Route53 assigns four NS records (e.g., ns-1234.awsdns-56.org)
3. Go to your registrar's dashboard
4. Replace the default nameservers with the four Route53 NS records
5. Wait for propagation (can take up to 48 hours, usually much faster)
6. Now Route53 is authoritative for yourdomain.com</code></pre><p>Once the delegation is complete, any DNS records you create in your Route53 hosted zone will be the ones the internet sees.</p><h5><strong>Route53 Alias records: a special AWS feature</strong></h5><p>Standard DNS has a limitation: you cannot put a CNAME record at the zone apex (the bare domain like <code>yourdomain.com</code>). This is a problem because AWS resources like ALBs and CloudFront distributions do not have static IP addresses, so you cannot use an A record either.</p><p>Route53 solves this with <strong>Alias records</strong>. An Alias record looks like an A or AAAA record to DNS clients, but behind the scenes it resolves to another AWS resource. Think of it as a CNAME that works at the zone apex.</p><blockquote><ul><li><p><strong>No charge</strong>: Route53 does not charge for queries to Alias records that point at AWS resources.</p></li><li><p><strong>Zone apex compatible</strong>: You can create an Alias record for <code>yourdomain.com</code> pointing at your ALB.</p></li><li><p><strong>Health check aware</strong>: Alias records can inherit health checks from the target resource.</p></li></ul></blockquote><p>We will use an Alias record to point our domain at our Application Load Balancer.</p><h5><strong>TLS/SSL: why HTTPS matters</strong></h5><p>Right now our ALB is serving traffic over HTTP on port 80. Every request and response travels across the internet in plain text. Anyone between the user and your server (ISPs, Wi-Fi operators, anyone on the same network) can read the data, modify it, or inject content.</p><p>TLS (Transport Layer Security) encrypts the connection between the user&#8217;s browser and your server. When you see the padlock icon and <code>https://</code> in your browser, that means TLS is in use.</p><p>Why HTTPS matters:</p><blockquote><ul><li><p><strong>Confidentiality</strong>: Data in transit is encrypted. Passwords, API keys, personal information are all protected.</p></li><li><p><strong>Integrity</strong>: Data cannot be modified in transit. No one can inject ads, malware, or tracking scripts into your responses.</p></li><li><p><strong>Authentication</strong>: The certificate proves that the server is who it claims to be. This prevents man-in-the-middle attacks.</p></li><li><p><strong>SEO and trust</strong>: Google has used HTTPS as a ranking signal since 2014. Browsers mark HTTP sites as &#8220;Not Secure.&#8221; Users trust HTTPS sites more.</p></li><li><p><strong>Required for modern features</strong>: HTTP/2, service workers, geolocation API, and many other browser features require HTTPS.</p></li></ul></blockquote><p>In short, there is no good reason to serve production traffic over HTTP.</p><h5><strong>How TLS works (the short version)</strong></h5><p>When your browser connects to an HTTPS server, a process called the TLS handshake happens before any application data is exchanged:</p><pre><code>Client                              Server
  |                                    |
  |--- ClientHello (supported ciphers) --&gt;|
  |                                    |
  |&lt;-- ServerHello + Certificate -------|
  |                                    |
  |--- Key exchange material ----------&gt;|
  |                                    |
  |&lt;-- Key exchange material -----------|
  |                                    |
  |   (Both sides derive session key)  |
  |                                    |
  |&lt;== Encrypted application data ====&gt;|</code></pre><blockquote><ul><li><p><strong>ClientHello</strong>: The browser sends the TLS versions and cipher suites it supports.</p></li><li><p><strong>ServerHello</strong>: The server picks a cipher suite and sends its certificate (which contains the server&#8217;s public key).</p></li><li><p><strong>Certificate validation</strong>: The browser checks that the certificate was issued by a trusted Certificate Authority (CA), is not expired, and matches the domain name.</p></li><li><p><strong>Key exchange</strong>: Both sides exchange key material and derive a shared session key.</p></li><li><p><strong>Encrypted communication</strong>: All subsequent data is encrypted with the session key using symmetric encryption (much faster than asymmetric).</p></li></ul></blockquote><p>The important takeaway is that you need a valid TLS certificate for your domain. That is where AWS Certificate Manager comes in.</p><h5><strong>AWS Certificate Manager (ACM)</strong></h5><p>ACM is a free service that lets you provision, manage, and deploy TLS certificates for use with AWS services like ALB, CloudFront, and API Gateway.</p><p>The key benefits:</p><blockquote><ul><li><p><strong>Free</strong>: Public certificates are free when used with AWS services. No need to pay a CA.</p></li><li><p><strong>Auto-renewal</strong>: ACM automatically renews certificates before they expire. No more 3 AM pages because a certificate expired.</p></li><li><p><strong>DNS validation</strong>: You prove domain ownership by adding a CNAME record to your DNS. This is fully automatable with Terraform.</p></li><li><p><strong>Managed private keys</strong>: ACM stores the private key securely. You never have to handle it yourself.</p></li></ul></blockquote><p>The process for getting a certificate with ACM looks like this:</p><pre><code>1. Request a certificate for api.yourdomain.com (and optionally *.yourdomain.com)
2. ACM gives you a CNAME record to add to your DNS
3. You add the CNAME record to your Route53 hosted zone
4. ACM validates that you own the domain
5. Certificate is issued (usually within minutes)
6. ACM auto-renews it every 13 months</code></pre><p>DNS validation is preferred over email validation because it can be fully automated and does not require someone to click a link in an email.</p><h5><strong>Connecting it all: the full picture</strong></h5><p>Let&#8217;s put together everything we have discussed. Here is how all the pieces connect to make your application reachable over HTTPS at a real domain:</p><pre><code>User's browser
     |
     | (DNS query: api.yourdomain.com)
     v
Route53 hosted zone
     |
     | (Alias record -&gt; ALB)
     v
Application Load Balancer
     |
     |--- Port 443 (HTTPS) -&gt; Forward to target group (with ACM certificate)
     |--- Port 80  (HTTP)  -&gt; Redirect to HTTPS
     |
     v
ECS Fargate tasks (your API containers)</code></pre><blockquote><ul><li><p><strong>Route53</strong> resolves <code>api.yourdomain.com</code> to the ALB&#8217;s address using an Alias record.</p></li><li><p><strong>ACM</strong> provides the TLS certificate that the ALB uses for HTTPS termination.</p></li><li><p><strong>ALB</strong> terminates TLS, meaning the encrypted connection ends at the ALB. Traffic between the ALB and your ECS tasks travels over HTTP within your VPC (which is fine because it is on a private network).</p></li><li><p><strong>HTTP to HTTPS redirect</strong>: The ALB listens on port 80 and automatically redirects to port 443, so users who type <code>http://</code> still end up on HTTPS.</p></li></ul></blockquote><h5><strong>Terraform: Route53 hosted zone</strong></h5><p>Let&#8217;s write the Terraform code. We will build on the infrastructure from article eight. First, the Route53 hosted zone:</p><pre><code># dns.tf
variable "domain_name" {
  description = "The root domain name"
  type        = string
  default     = "yourdomain.com"
}

variable "api_subdomain" {
  description = "Subdomain for the API"
  type        = string
  default     = "api"
}

resource "aws_route53_zone" "main" {
  name = var.domain_name

  tags = {
    Name = var.domain_name
  }
}</code></pre><p>This creates a public hosted zone. Route53 automatically creates the NS and SOA records. After applying this, you need to copy the four NS records and configure them at your domain registrar. You only need to do this once.</p><p>You can output the nameservers so you know what to set:</p><pre><code>output "nameservers" {
  description = "Nameservers for the hosted zone. Set these at your registrar."
  value       = aws_route53_zone.main.name_servers
}</code></pre><h5><strong>Terraform: ACM certificate with DNS validation</strong></h5><p>Next, we request a TLS certificate and validate it automatically through DNS:</p><pre><code># acm.tf
resource "aws_acm_certificate" "app" {
  domain_name               = "${var.api_subdomain}.${var.domain_name}"
  subject_alternative_names = ["*.${var.domain_name}"]
  validation_method         = "DNS"

  lifecycle {
    create_before_destroy = true
  }

  tags = {
    Name = "${var.api_subdomain}.${var.domain_name}"
  }
}

resource "aws_route53_record" "acm_validation" {
  for_each = {
    for dvo in aws_acm_certificate.app.domain_validation_options : dvo.domain_name =&gt; {
      name   = dvo.resource_record_name
      record = dvo.resource_record_value
      type   = dvo.resource_record_type
    }
  }

  allow_overwrite = true
  name            = each.value.name
  records         = [each.value.record]
  ttl             = 60
  type            = each.value.type
  zone_id         = aws_route53_zone.main.zone_id
}

resource "aws_acm_certificate_validation" "app" {
  certificate_arn         = aws_acm_certificate.app.arn
  validation_record_fqdns = [for record in aws_route53_record.acm_validation : record.fqdn]
}</code></pre><p>Let&#8217;s break this down:</p><blockquote><ul><li><p><code>aws_acm_certificate</code> requests the certificate. We request it for <code>api.yourdomain.com</code> with a wildcard SAN (<code>*.yourdomain.com</code>) so it covers any subdomain.</p></li><li><p><code>validation_method = "DNS"</code> tells ACM we will prove domain ownership by adding DNS records.</p></li><li><p><code>create_before_destroy = true</code> ensures that when renewing, the new certificate is created before the old one is destroyed. This prevents downtime.</p></li><li><p><code>aws_route53_record.acm_validation</code> creates the CNAME records that ACM requires for validation. The <code>for_each</code> loop handles the case where the certificate covers multiple domain names.</p></li><li><p><code>aws_acm_certificate_validation</code> is a waiter resource. Terraform will block here until ACM confirms the certificate is validated and issued. This usually takes 2-5 minutes.</p></li></ul></blockquote><h5><strong>Terraform: ALB HTTPS listener and HTTP redirect</strong></h5><p>In article eight, we created an ALB with only an HTTP listener. Now we are going to add an HTTPS listener and change the HTTP listener to redirect to HTTPS:</p><pre><code># alb.tf (updated)

# Change the existing HTTP listener to redirect
resource "aws_lb_listener" "http" {
  load_balancer_arn = aws_lb.app.arn
  port              = 80
  protocol          = "HTTP"

  default_action {
    type = "redirect"

    redirect {
      port        = "443"
      protocol    = "HTTPS"
      status_code = "HTTP_301"
    }
  }
}

# Add HTTPS listener
resource "aws_lb_listener" "https" {
  load_balancer_arn = aws_lb.app.arn
  port              = 443
  protocol          = "HTTPS"
  ssl_policy        = "ELBSecurityPolicy-TLS13-1-2-2021-06"
  certificate_arn   = aws_acm_certificate_validation.app.certificate_arn

  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.app.arn
  }
}</code></pre><p>Important details:</p><blockquote><ul><li><p>The HTTP listener now returns a <strong>301 redirect</strong> to HTTPS. This is a permanent redirect, so browsers and search engines will remember it and go directly to HTTPS next time.</p></li><li><p>The HTTPS listener references the validated ACM certificate. Note that we reference <code>aws_acm_certificate_validation.app.certificate_arn</code>, not the certificate directly. This ensures Terraform waits for validation to complete before creating the listener.</p></li><li><p><code>ssl_policy</code> controls which TLS versions and cipher suites the ALB accepts. <code>ELBSecurityPolicy-TLS13-1-2-2021-06</code> supports TLS 1.2 and 1.3, which is the current best practice. Older policies that allow TLS 1.0 or 1.1 should not be used.</p></li></ul></blockquote><p>You also need to update the ALB security group to allow HTTPS traffic:</p><pre><code># security_groups.tf (updated)
resource "aws_security_group" "alb" {
  name        = "${var.project_name}-alb-sg"
  description = "Security group for the Application Load Balancer"
  vpc_id      = aws_vpc.main.id

  ingress {
    description = "HTTP from anywhere (for redirect)"
    from_port   = 80
    to_port     = 80
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  ingress {
    description = "HTTPS from anywhere"
    from_port   = 443
    to_port     = 443
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }

  tags = {
    Name = "${var.project_name}-alb-sg"
  }
}</code></pre><p>We keep port 80 open so that the redirect works. If you close port 80, users who type <code>http://</code> will get a connection timeout instead of a redirect.</p><h5><strong>Terraform: Route53 record pointing to the ALB</strong></h5><p>Finally, we create the DNS record that points our domain at the load balancer:</p><pre><code># dns.tf (continued)
resource "aws_route53_record" "api" {
  zone_id = aws_route53_zone.main.zone_id
  name    = "${var.api_subdomain}.${var.domain_name}"
  type    = "A"

  alias {
    name                   = aws_lb.app.dns_name
    zone_id                = aws_lb.app.zone_id
    evaluate_target_health = true
  }
}</code></pre><p>This is an Alias record. Even though it is <code>type = "A"</code>, it does not contain a hardcoded IP address. Instead, it points to the ALB&#8217;s DNS name. Route53 resolves the ALB&#8217;s current IP addresses behind the scenes and returns them to the client.</p><p><code>evaluate_target_health = true</code> means that if the ALB has no healthy targets, Route53 will not return this record in DNS queries. This is useful in multi-region setups with failover routing.</p><h5><strong>Health checks</strong></h5><p>Health checks are how AWS determines whether your application is actually working. There are two levels of health checks in our setup:</p><p><strong>ALB target group health checks</strong></p><p>We already configured these in article eight. The ALB periodically sends HTTP requests to your ECS tasks on a path you define (like <code>/health</code>). If a task fails consecutive checks, the ALB stops sending it traffic and ECS replaces it.</p><pre><code># Already in our target group from article 8
health_check {
  enabled             = true
  healthy_threshold   = 3
  unhealthy_threshold = 3
  timeout             = 5
  interval            = 30
  path                = "/health"
  protocol            = "HTTP"
  matcher             = "200"
}</code></pre><p><strong>Route53 health checks</strong></p><p>Route53 health checks operate at the DNS level. They monitor an endpoint and can trigger DNS failover if the endpoint goes down. These are particularly useful when you have resources in multiple regions:</p><pre><code># route53_health.tf
resource "aws_route53_health_check" "api" {
  fqdn              = "${var.api_subdomain}.${var.domain_name}"
  port               = 443
  type               = "HTTPS"
  resource_path      = "/health"
  failure_threshold  = 3
  request_interval   = 30
  measure_latency    = true

  tags = {
    Name = "${var.api_subdomain}.${var.domain_name}-health-check"
  }
}</code></pre><blockquote><ul><li><p><code>type = "HTTPS"</code> means Route53 connects over TLS to check the endpoint.</p></li><li><p><code>failure_threshold = 3</code> means three consecutive failures mark the endpoint as unhealthy.</p></li><li><p><code>request_interval = 30</code> checks every 30 seconds. You can set this to 10 for faster detection, but it costs more.</p></li><li><p><code>measure_latency = true</code> tracks latency metrics in CloudWatch.</p></li></ul></blockquote><p>For a single-region setup, Route53 health checks are optional but nice to have for monitoring. For multi-region with failover routing, they are essential.</p><h5><strong>The full Terraform configuration</strong></h5><p>Let&#8217;s put it all together in one view so you can see how the pieces connect:</p><pre><code># Full dns.tf
variable "domain_name" {
  description = "The root domain name"
  type        = string
  default     = "yourdomain.com"
}

variable "api_subdomain" {
  description = "Subdomain for the API"
  type        = string
  default     = "api"
}

# Hosted zone
resource "aws_route53_zone" "main" {
  name = var.domain_name

  tags = {
    Name = var.domain_name
  }
}

# ACM certificate
resource "aws_acm_certificate" "app" {
  domain_name               = "${var.api_subdomain}.${var.domain_name}"
  subject_alternative_names = ["*.${var.domain_name}"]
  validation_method         = "DNS"

  lifecycle {
    create_before_destroy = true
  }

  tags = {
    Name = "${var.api_subdomain}.${var.domain_name}"
  }
}

# DNS validation records for ACM
resource "aws_route53_record" "acm_validation" {
  for_each = {
    for dvo in aws_acm_certificate.app.domain_validation_options : dvo.domain_name =&gt; {
      name   = dvo.resource_record_name
      record = dvo.resource_record_value
      type   = dvo.resource_record_type
    }
  }

  allow_overwrite = true
  name            = each.value.name
  records         = [each.value.record]
  ttl             = 60
  type            = each.value.type
  zone_id         = aws_route53_zone.main.zone_id
}

# Wait for certificate validation
resource "aws_acm_certificate_validation" "app" {
  certificate_arn         = aws_acm_certificate.app.arn
  validation_record_fqdns = [for record in aws_route53_record.acm_validation : record.fqdn]
}

# DNS record pointing to ALB
resource "aws_route53_record" "api" {
  zone_id = aws_route53_zone.main.zone_id
  name    = "${var.api_subdomain}.${var.domain_name}"
  type    = "A"

  alias {
    name                   = aws_lb.app.dns_name
    zone_id                = aws_lb.app.zone_id
    evaluate_target_health = true
  }
}

# Outputs
output "nameservers" {
  description = "Nameservers for the hosted zone. Set these at your registrar."
  value       = aws_route53_zone.main.name_servers
}

output "app_url" {
  description = "The HTTPS URL for the API"
  value       = "https://${var.api_subdomain}.${var.domain_name}"
}</code></pre><h5><strong>Applying the changes</strong></h5><p>Run <code>terraform plan</code> first to see what will be created:</p><pre><code>terraform plan</code></pre><p>You should see new resources for the hosted zone, ACM certificate, validation records, HTTPS listener, and the DNS record. Once you are satisfied with the plan:</p><pre><code>terraform apply</code></pre><p>The <code>aws_acm_certificate_validation</code> resource will block until the certificate is validated. This usually takes 2-5 minutes. If it takes longer than 10 minutes, check that the validation CNAME records were created correctly in the hosted zone and that your nameservers are properly delegated.</p><p>After the apply completes, update your nameservers at your registrar if you have not already. Then verify everything works:</p><pre><code># Check DNS resolution
dig api.yourdomain.com

# Test HTTPS
curl -v https://api.yourdomain.com/health

# Test HTTP redirect
curl -v http://api.yourdomain.com/health
# Should return a 301 redirect to https://</code></pre><h5><strong>Debugging DNS issues</strong></h5><p>DNS problems are some of the most frustrating to debug because of caching. Here are the tools and techniques you need:</p><pre><code># Query a specific nameserver directly (bypass cache)
dig @ns-1234.awsdns-56.org api.yourdomain.com

# Check all record types
dig api.yourdomain.com ANY

# Trace the full resolution path
dig +trace api.yourdomain.com

# Check nameserver delegation
dig yourdomain.com NS

# Check TXT records (useful for ACM validation)
dig _acme-challenge.api.yourdomain.com TXT</code></pre><p>Common issues and how to fix them:</p><blockquote><ul><li><p><strong>&#8220;NXDOMAIN&#8221; (domain not found)</strong>: Your nameservers are not delegated correctly. Check the NS records at your registrar.</p></li><li><p><strong>Old IP address returned</strong>: DNS caching. Wait for the TTL to expire, or use <code>dig @8.8.8.8</code> to check Google&#8217;s resolvers directly.</p></li><li><p><strong>ACM validation stuck</strong>: The CNAME record name and value must match exactly what ACM expects. Check for trailing dots or typos.</p></li><li><p><strong>Certificate not valid for domain</strong>: The certificate&#8217;s Common Name or SAN does not match the domain. Make sure you requested the certificate for the correct domain name.</p></li></ul></blockquote><h5><strong>A note about CloudFront (CDN)</strong></h5><p>So far we have connected users directly to our ALB through Route53. This works well, but for applications that serve static assets (images, CSS, JavaScript) or have users spread across the globe, you should consider putting CloudFront in front of your ALB.</p><p>CloudFront is AWS&#8217;s CDN (Content Delivery Network). It caches your content at edge locations around the world, so users get responses from a server that is geographically close to them instead of from your origin region.</p><pre><code>Without CloudFront:
  User in Tokyo -&gt; Route53 -&gt; ALB in us-east-1 (200ms latency)

With CloudFront:
  User in Tokyo -&gt; Route53 -&gt; CloudFront edge in Tokyo (cached, 20ms)
                                    |
                                    v (cache miss only)
                              ALB in us-east-1</code></pre><p>Benefits of CloudFront:</p><blockquote><ul><li><p><strong>Lower latency</strong>: Content served from the nearest edge location.</p></li><li><p><strong>Reduced origin load</strong>: Cached responses do not hit your ALB or ECS tasks.</p></li><li><p><strong>DDoS protection</strong>: CloudFront integrates with AWS Shield for DDoS mitigation.</p></li><li><p><strong>Free ACM certificates</strong>: CloudFront uses ACM certificates from us-east-1 (this is a requirement, the certificate must be in us-east-1 regardless of where your origin is).</p></li></ul></blockquote><p>We will not set up CloudFront in this article since it deserves its own deep dive, but keep it in mind for when you need to optimize performance for a global audience.</p><h5><strong>Cost breakdown</strong></h5><p>Let&#8217;s look at what this setup costs:</p><blockquote><ul><li><p><strong>Route53 hosted zone</strong>: $0.50/month per hosted zone.</p></li><li><p><strong>Route53 queries</strong>: $0.40 per million queries. For most applications this is negligible.</p></li><li><p><strong>Route53 health checks</strong>: $0.50/month for a basic HTTPS health check. $1.00/month with latency measurement.</p></li><li><p><strong>ACM certificates</strong>: Free when used with AWS services (ALB, CloudFront, API Gateway).</p></li><li><p><strong>ALB</strong>: The ALB was already part of our ECS setup. No additional cost for HTTPS termination.</p></li></ul></blockquote><p>Total additional cost for DNS and TLS: roughly $1-2/month. This is one of the cheapest and highest value improvements you can make to your infrastructure.</p><h5><strong>Security best practices</strong></h5><p>Before we wrap up, here are some security practices to keep in mind:</p><blockquote><ul><li><p><strong>Always redirect HTTP to HTTPS</strong>: Never serve production traffic over plain HTTP.</p></li><li><p><strong>Use a modern TLS policy</strong>: <code>ELBSecurityPolicy-TLS13-1-2-2021-06</code> or newer. Disable TLS 1.0 and 1.1.</p></li><li><p><strong>Enable HSTS</strong>: Add the <code>Strict-Transport-Security</code> header in your application to tell browsers to always use HTTPS. This prevents downgrade attacks.</p></li><li><p><strong>Use separate certificates per environment</strong>: Do not reuse production certificates in staging. ACM is free, so there is no reason not to have separate certificates.</p></li><li><p><strong>Monitor certificate expiry</strong>: Even though ACM auto-renews, set up a CloudWatch alarm for certificate expiry as a safety net. If DNS validation fails for some reason, auto-renewal will fail silently.</p></li></ul></blockquote><h5><strong>Closing notes</strong></h5><p>Your application is now reachable at a real domain over HTTPS. We covered a lot of ground in this article: DNS fundamentals and record types, Route53 hosted zones and routing policies, TLS certificates with ACM, HTTPS termination at the ALB, HTTP to HTTPS redirects, health checks at both the ALB and DNS level, and the Terraform code to provision all of it.</p><p>This is a milestone in the series. Your TypeScript API is now running in containers on ECS, behind a load balancer, with auto-scaling, accessible at a clean URL over an encrypted connection. That is a production-grade setup.</p><p>In the next article, we will look at monitoring and observability so you can see what your application is doing in production and catch problems before your users do.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Secrets, Config, and Environment Management]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-secrets-and-config</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-secrets-and-config</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Fri, 15 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JqxA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JqxA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JqxA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!JqxA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!JqxA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!JqxA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JqxA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043138?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JqxA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!JqxA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!JqxA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!JqxA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76363c93-54aa-4b43-bb06-2392a0e4bef6_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article nine of the DevOps from Zero to Hero series. In the previous article we deployed our TypeScript API to ECS with Fargate, and everything is running in the cloud. But we skipped over something important: how does your application get its database URL, API keys, and other configuration values? If you hard-coded them into your source code, you have a problem.</p><p>Configuration and secrets management is one of those topics that seems simple until you get it wrong. A leaked API key can cost you thousands of dollars. A misconfigured database URL can point your production app at the staging database. A checked-in <code>.env</code> file can expose credentials to anyone who clones your repository. These are not hypothetical scenarios, they happen all the time.</p><p>In this article we will cover the foundational practices for handling configuration and secrets: the 12-factor methodology, environment variables, <code>.env</code> files, secret scanning, AWS Secrets Manager, AWS Systems Manager Parameter Store, and how to structure configuration across dev, staging, and production environments. By the end you will have a clear, practical approach to keeping your config clean and your secrets safe.</p><p>Let&#8217;s get into it.</p><h5><strong>The 12-factor app: config belongs in the environment</strong></h5><p>The <a href="https://12factor.net/">Twelve-Factor App</a> is a methodology for building modern applications that was published by the team at Heroku back in 2012. It describes twelve principles for building software that is easy to deploy, scale, and maintain. Factor number three is about configuration, and it says something very clear: store config in the environment.</p><p>What does &#8220;config&#8221; mean here? It is anything that is likely to change between environments (dev, staging, production). Database URLs, API keys, feature flags, external service endpoints, log levels. These values should not live in your source code. They should not be baked into your Docker image. They should come from the environment where your application is running.</p><p>The reasoning is simple:</p><blockquote><ul><li><p><strong>Security</strong>: Secrets in source code end up in version control, in CI logs, in Docker layers, and in the hands of anyone who has access to your repository.</p></li><li><p><strong>Portability</strong>: If your database URL is hard-coded, you cannot run the same code against a staging database without changing the code. If it comes from the environment, you just change the environment variable.</p></li><li><p><strong>Simplicity</strong>: One build artifact (your Docker image) works in every environment. The only thing that changes is the configuration injected at runtime.</p></li></ul></blockquote><p>Here is the anti-pattern versus the correct approach:</p><pre><code>// BAD: hard-coded config
const dbUrl = "postgresql://admin:supersecret@prod-db.example.com:5432/myapp";

// GOOD: read from the environment
const dbUrl = process.env.DATABASE_URL;
if (!dbUrl) {
  throw new Error("DATABASE_URL environment variable is required");
}</code></pre><p>That second example follows the 12-factor principle. The application does not know or care which environment it is running in. It just reads the value from the environment and uses it.</p><h5><strong>Environment variables: how they work</strong></h5><p>Environment variables are key-value pairs that exist in the operating system&#8217;s process environment. Every process inherits the environment of its parent process, and you can set additional variables when launching a process.</p><p>Setting and reading environment variables in the shell:</p><pre><code># Set a variable for the current shell session
export DATABASE_URL="postgresql://localhost:5432/myapp"

# Read it
echo $DATABASE_URL

# Set a variable only for a single command
DATABASE_URL="postgresql://localhost:5432/myapp" node app.js

# List all environment variables
env

# Unset a variable
unset DATABASE_URL</code></pre><p>In Node.js/TypeScript, you access them through <code>process.env</code>:</p><pre><code>// Read an environment variable
const port = process.env.PORT || "3000";
const dbUrl = process.env.DATABASE_URL;
const logLevel = process.env.LOG_LEVEL || "info";

// Check for required variables at startup
const required = ["DATABASE_URL", "API_KEY", "JWT_SECRET"];
for (const key of required) {
  if (!process.env[key]) {
    console.error(`Missing required environment variable: ${key}`);
    process.exit(1);
  }
}</code></pre><p>This pattern of checking for required variables at startup is important. You want your application to fail fast and loud if it is missing configuration, not silently break at some random point later.</p><h5><strong>Dotenv files: local development convenience</strong></h5><p>Typing <code>export DATABASE_URL=...</code> every time you open a terminal gets old fast. That is where <code>.env</code> files come in. A <code>.env</code> file is a simple text file that lists environment variables, one per line:</p><pre><code># .env
DATABASE_URL=postgresql://localhost:5432/myapp_dev
API_KEY=dev-api-key-not-real
JWT_SECRET=local-dev-secret
LOG_LEVEL=debug
PORT=3000</code></pre><p>Libraries like <a href="https://www.npmjs.com/package/dotenv">dotenv</a> for Node.js automatically read this file and load the variables into <code>process.env</code> when your application starts:</p><pre><code>// Load .env file at the very top of your entry point
import "dotenv/config";

// Now process.env.DATABASE_URL is available
console.log(process.env.DATABASE_URL);</code></pre><p>The critical rule with <code>.env</code> files is: <strong>never commit them to Git</strong>. They contain secrets, and your Git repository is not a secure place to store secrets. Add <code>.env</code> to your <code>.gitignore</code> immediately:</p><pre><code># .gitignore

# Environment files with secrets
.env
.env.local
.env.*.local

# Keep the example file (no real secrets)
!.env.example</code></pre><p>Instead of committing your actual <code>.env</code> file, commit a <code>.env.example</code> file with placeholder values. This tells your teammates what variables they need without exposing real secrets:</p><pre><code># .env.example
DATABASE_URL=postgresql://localhost:5432/myapp_dev
API_KEY=your-api-key-here
JWT_SECRET=generate-a-random-string
LOG_LEVEL=debug
PORT=3000</code></pre><p>When a new developer joins the team, they copy <code>.env.example</code> to <code>.env</code> and fill in their own values. Simple, safe, effective.</p><h5><strong>Why you should never commit secrets to Git</strong></h5><p>This deserves its own section because it is that important. When you commit a secret to Git, it does not just exist in the current version of the file. It exists in the Git history forever. Even if you delete the file or overwrite the value in a later commit, anyone who clones the repository can find it by looking at the commit history.</p><pre><code># Oops, I committed my .env file
git log --all --full-history -- .env

# Anyone can see the contents of that file at that commit
git show abc123:.env</code></pre><p>If this happens, the secret is compromised. You need to rotate it immediately, meaning generate a new key and revoke the old one. Rewriting Git history with <code>git filter-branch</code> or BFG Repo-Cleaner is possible but painful, especially in a shared repository.</p><p>The better approach is prevention. Use tools that scan your repository for secrets before they ever get committed:</p><blockquote><ul><li><p><strong><a href="https://github.com/awslabs/git-secrets">git-secrets</a></strong>: An AWS tool that installs Git hooks to prevent committing secrets. It scans for AWS access keys, secret keys, and custom patterns you define.</p></li><li><p><strong><a href="https://github.com/gitleaks/gitleaks">gitleaks</a></strong>: A faster, more comprehensive scanner that detects a wide range of secret patterns (API keys, tokens, passwords) across your entire repository history.</p></li><li><p><strong><a href="https://pre-commit.com/">pre-commit</a></strong>: A framework for managing Git pre-commit hooks. You can add gitleaks or git-secrets as a hook that runs automatically on every commit.</p></li></ul></blockquote><p>Here is how to set up gitleaks as a pre-commit hook:</p><pre><code># .pre-commit-config.yaml
repos:
  - repo: https://github.com/gitleaks/gitleaks
    rev: v8.18.0
    hooks:
      - id: gitleaks</code></pre><pre><code># Install pre-commit and set up the hooks
pip install pre-commit
pre-commit install

# Now every commit will be scanned for secrets automatically
git commit -m "add new feature"
# gitleaks runs and blocks the commit if it finds a secret</code></pre><p>You should also run gitleaks in your CI pipeline as a safety net. We covered CI pipelines in article five, so adding a gitleaks step is straightforward:</p><pre><code># In your GitHub Actions workflow
- name: Scan for secrets
  uses: gitleaks/gitleaks-action@v2
  env:
    GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}</code></pre><h5><strong>Config hierarchy: how values get resolved</strong></h5><p>In a real application, configuration can come from multiple sources. When the same key is defined in more than one place, you need a clear precedence order. The standard hierarchy, from lowest to highest priority, looks like this:</p><pre><code>1. Application defaults (hard-coded fallbacks in your code)
2. Config files (JSON, YAML, TOML files loaded at startup)
3. Environment variables (set by the OS, container runtime, or .env file)
4. CLI flags (passed when starting the application)
5. Remote config (fetched from Secrets Manager, Parameter Store, etc.)</code></pre><p>Each level overrides the one below it. So if your code has a default <code>LOG_LEVEL=info</code>, your config file sets it to <code>warn</code>, and your environment variable sets it to <code>debug</code>, the environment variable wins. If you also pass <code>--log-level=error</code> as a CLI flag, that wins over everything.</p><p>Here is a practical example showing this hierarchy in TypeScript:</p><pre><code>import { readFileSync, existsSync } from "fs";

interface AppConfig {
  port: number;
  logLevel: string;
  dbUrl: string;
}

function loadConfig(): AppConfig {
  // Level 1: Application defaults
  let config: AppConfig = {
    port: 3000,
    logLevel: "info",
    dbUrl: "postgresql://localhost:5432/myapp",
  };

  // Level 2: Config file (if it exists)
  const configPath = "./config.json";
  if (existsSync(configPath)) {
    const fileConfig = JSON.parse(readFileSync(configPath, "utf-8"));
    config = { ...config, ...fileConfig };
  }

  // Level 3: Environment variables (override file config)
  if (process.env.PORT) config.port = parseInt(process.env.PORT, 10);
  if (process.env.LOG_LEVEL) config.logLevel = process.env.LOG_LEVEL;
  if (process.env.DATABASE_URL) config.dbUrl = process.env.DATABASE_URL;

  return config;
}

const config = loadConfig();
console.log("Config loaded:", config);</code></pre><p>This pattern gives you flexibility. Developers can use a config file locally, the CI environment can set environment variables, and production can pull secrets from AWS Secrets Manager (which we will cover next).</p><h5><strong>AWS Secrets Manager: storing and retrieving secrets</strong></h5><p>AWS Secrets Manager is a managed service for storing, retrieving, and rotating secrets. Unlike environment variables, which are visible in ECS task definitions, CloudFormation templates, and potentially in logs, Secrets Manager stores values encrypted at rest and provides fine-grained access control through IAM policies.</p><p>When should you use Secrets Manager instead of plain environment variables?</p><blockquote><ul><li><p><strong>Database credentials</strong>: Secrets Manager can automatically rotate database passwords on a schedule, updating both the secret value and the database itself.</p></li><li><p><strong>API keys for third-party services</strong>: Stripe, Twilio, SendGrid, anything where a leaked key means real money.</p></li><li><p><strong>TLS certificates and private keys</strong>: Anything cryptographic that should never appear in plain text.</p></li><li><p><strong>Shared secrets across services</strong>: When multiple services need the same credentials, Secrets Manager is a single source of truth.</p></li></ul></blockquote><p>Creating a secret with the AWS CLI:</p><pre><code># Create a simple string secret
aws secretsmanager create-secret \
  --name "prod/task-api/database-url" \
  --description "Production database connection string" \
  --secret-string "postgresql://admin:s3cur3P@ss@prod-db.example.com:5432/myapp"

# Create a JSON secret (multiple key-value pairs in one secret)
aws secretsmanager create-secret \
  --name "prod/task-api/credentials" \
  --description "Production API credentials" \
  --secret-string '{
    "DB_URL": "postgresql://admin:s3cur3P@ss@prod-db.example.com:5432/myapp",
    "API_KEY": "sk_live_abc123",
    "JWT_SECRET": "a-very-long-random-string"
  }'</code></pre><p>Notice the naming convention: <code>environment/service/secret-name</code>. This hierarchical naming makes it easy to organize secrets and write IAM policies that restrict access by environment or service.</p><p>Retrieving a secret:</p><pre><code># Get the secret value
aws secretsmanager get-secret-value \
  --secret-id "prod/task-api/database-url" \
  --query SecretString \
  --output text</code></pre><h5><strong>Secrets Manager: rotation basics</strong></h5><p>One of the most powerful features of Secrets Manager is automatic rotation. Instead of using the same database password forever (and hoping nobody leaks it), you can configure Secrets Manager to rotate the password on a schedule, for example every 30 days.</p><p>For Amazon RDS databases, AWS provides built-in rotation Lambda functions. The rotation process works like this:</p><pre><code>1. Secrets Manager invokes a Lambda function on a schedule
2. The Lambda generates a new password
3. It updates the password in the RDS database
4. It stores the new password in Secrets Manager
5. Your application fetches the new value next time it reads the secret</code></pre><p>Setting up rotation with the CLI:</p><pre><code># Enable rotation for an RDS secret
aws secretsmanager rotate-secret \
  --secret-id "prod/task-api/database-url" \
  --rotation-lambda-arn "arn:aws:lambda:us-east-1:123456789012:function:SecretsManagerRDSRotation" \
  --rotation-rules '{"AutomaticallyAfterDays": 30}'</code></pre><p>The important thing to understand about rotation is that your application needs to handle it gracefully. If your app caches the database connection string at startup and never re-reads it, a rotated password will break your connection. The solution is to either re-fetch the secret periodically or use a connection library that can handle credential refresh.</p><h5><strong>Secrets Manager: IAM access policies</strong></h5><p>You control who and what can access your secrets through IAM policies. Here is a policy that allows an ECS task role to read only the secrets for a specific environment and service:</p><pre><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "secretsmanager:GetSecretValue",
        "secretsmanager:DescribeSecret"
      ],
      "Resource": "arn:aws:secretsmanager:us-east-1:123456789012:secret:prod/task-api/*"
    }
  ]
}</code></pre><p>This policy follows the principle of least privilege. The ECS task can only read secrets under the <code>prod/task-api/</code> prefix. It cannot list all secrets in the account, it cannot read secrets from other services, and it cannot modify or delete any secrets. If someone compromises your task-api service, they still cannot access the secrets belonging to your user-service or payment-service.</p><p>You attach this policy to the ECS task execution role that we set up in the previous article:</p><pre><code># Create the policy
aws iam create-policy \
  --policy-name task-api-secrets-read \
  --policy-document file://secrets-policy.json

# Attach it to the ECS task role
aws iam attach-role-policy \
  --role-name task-api-task-role \
  --policy-arn "arn:aws:iam::123456789012:policy/task-api-secrets-read"</code></pre><h5><strong>AWS Systems Manager Parameter Store</strong></h5><p>Parameter Store is another AWS service for storing configuration, and it serves a different purpose than Secrets Manager. Think of it this way:</p><blockquote><ul><li><p><strong>Secrets Manager</strong>: For sensitive values that need encryption, rotation, and fine-grained access control. It costs $0.40 per secret per month.</p></li><li><p><strong>Parameter Store</strong>: For non-sensitive or less-sensitive configuration values. The standard tier is free for up to 10,000 parameters.</p></li></ul></blockquote><p>Parameter Store supports three types of parameters:</p><blockquote><ul><li><p><strong>String</strong>: A plain text value. Good for configuration like log levels, feature flags, or endpoint URLs.</p></li><li><p><strong>StringList</strong>: A comma-separated list of values.</p></li><li><p><strong>SecureString</strong>: An encrypted value using AWS KMS. This provides similar encryption to Secrets Manager but without the rotation features.</p></li></ul></blockquote><p>Creating parameters with the CLI:</p><pre><code># Plain string parameter
aws ssm put-parameter \
  --name "/prod/task-api/log-level" \
  --type "String" \
  --value "info"

# Encrypted parameter
aws ssm put-parameter \
  --name "/prod/task-api/api-key" \
  --type "SecureString" \
  --value "sk_live_abc123"

# Get a parameter
aws ssm get-parameter \
  --name "/prod/task-api/log-level" \
  --query "Parameter.Value" \
  --output text

# Get an encrypted parameter (decrypt it)
aws ssm get-parameter \
  --name "/prod/task-api/api-key" \
  --with-decryption \
  --query "Parameter.Value" \
  --output text

# Get all parameters under a path
aws ssm get-parameters-by-path \
  --path "/prod/task-api/" \
  --with-decryption</code></pre><p>The hierarchical path naming (<code>/environment/service/parameter</code>) is the same convention we used with Secrets Manager, and it makes IAM policies straightforward:</p><pre><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": [
        "ssm:GetParameter",
        "ssm:GetParameters",
        "ssm:GetParametersByPath"
      ],
      "Resource": "arn:aws:ssm:us-east-1:123456789012:parameter/prod/task-api/*"
    }
  ]
}</code></pre><p>A common pattern is to use Parameter Store for non-sensitive config (log level, feature flags, service URLs) and Secrets Manager for truly sensitive values (database passwords, API keys). This keeps costs down and gives you the best of both services.</p><h5><strong>Practical example: loading config from env vars and Secrets Manager</strong></h5><p>Let&#8217;s bring everything together. Here is a realistic example of loading configuration in a TypeScript application that reads from environment variables first, then falls back to AWS Secrets Manager for sensitive values.</p><p>First, install the AWS SDK:</p><pre><code>npm install @aws-sdk/client-secrets-manager @aws-sdk/client-ssm</code></pre><p>Now the configuration loader:</p><pre><code>// src/config.ts
import {
  SecretsManagerClient,
  GetSecretValueCommand,
} from "@aws-sdk/client-secrets-manager";
import { SSMClient, GetParameterCommand } from "@aws-sdk/client-ssm";

const smClient = new SecretsManagerClient({ region: "us-east-1" });
const ssmClient = new SSMClient({ region: "us-east-1" });

interface AppConfig {
  port: number;
  logLevel: string;
  dbUrl: string;
  apiKey: string;
  jwtSecret: string;
}

async function getSecret(secretId: string): Promise&lt;string&gt; {
  const command = new GetSecretValueCommand({ SecretId: secretId });
  const response = await smClient.send(command);
  if (!response.SecretString) {
    throw new Error(`Secret ${secretId} has no string value`);
  }
  return response.SecretString;
}

async function getParameter(name: string): Promise&lt;string&gt; {
  const command = new GetParameterCommand({
    Name: name,
    WithDecryption: true,
  });
  const response = await ssmClient.send(command);
  if (!response.Parameter?.Value) {
    throw new Error(`Parameter ${name} not found`);
  }
  return response.Parameter.Value;
}

export async function loadConfig(): Promise&lt;AppConfig&gt; {
  const env = process.env.APP_ENV || "dev";

  // Non-sensitive config: prefer env vars, fall back to Parameter Store
  const port = process.env.PORT
    ? parseInt(process.env.PORT, 10)
    : 3000;

  const logLevel = process.env.LOG_LEVEL
    || await getParameter(`/${env}/task-api/log-level`).catch(() =&gt; "info");

  // Sensitive config: prefer env vars (for local dev), fall back to Secrets Manager
  let dbUrl = process.env.DATABASE_URL;
  let apiKey = process.env.API_KEY;
  let jwtSecret = process.env.JWT_SECRET;

  if (!dbUrl || !apiKey || !jwtSecret) {
    console.log(`Fetching secrets from AWS Secrets Manager for env: ${env}`);
    const secretString = await getSecret(`${env}/task-api/credentials`);
    const secrets = JSON.parse(secretString);

    dbUrl = dbUrl || secrets.DB_URL;
    apiKey = apiKey || secrets.API_KEY;
    jwtSecret = jwtSecret || secrets.JWT_SECRET;
  }

  if (!dbUrl || !apiKey || !jwtSecret) {
    throw new Error("Missing required configuration. Check env vars or Secrets Manager.");
  }

  return { port, logLevel, dbUrl, apiKey, jwtSecret };
}</code></pre><p>And here is how you use it in your application entry point:</p><pre><code>// src/index.ts
import "dotenv/config";
import { loadConfig } from "./config";
import { createApp } from "./app";

async function main() {
  const config = await loadConfig();
  console.log(`Starting server on port ${config.port} (log level: ${config.logLevel})`);

  const app = createApp(config);
  app.listen(config.port, () =&gt; {
    console.log(`Server running at http://localhost:${config.port}`);
  });
}

main().catch((err) =&gt; {
  console.error("Failed to start:", err);
  process.exit(1);
});</code></pre><p>This setup works for both local development and production:</p><blockquote><ul><li><p><strong>Local development</strong>: Developers set values in their <code>.env</code> file. The app reads from <code>process.env</code> and never hits AWS.</p></li><li><p><strong>Production</strong>: The <code>.env</code> file does not exist. The app detects the missing env vars and fetches from Secrets Manager. The ECS task role provides the necessary IAM permissions.</p></li></ul></blockquote><h5><strong>Environment promotion: dev, staging, and production</strong></h5><p>When you have multiple environments, you need a clear strategy for what changes between them and what stays the same. The general principle is: your code and Docker image should be identical across all environments. Only the configuration should differ.</p><p>Things that should differ between environments:</p><blockquote><ul><li><p><strong>Database connection strings</strong>: Each environment has its own database.</p></li><li><p><strong>API keys and secrets</strong>: Separate keys for each environment, so a compromised dev key does not affect production.</p></li><li><p><strong>Log levels</strong>: Usually <code>debug</code> in dev, <code>info</code> in staging, <code>warn</code> or <code>error</code> in production.</p></li><li><p><strong>Feature flags</strong>: Test new features in staging before enabling them in production.</p></li><li><p><strong>Scaling parameters</strong>: Dev runs one instance, production runs three or more.</p></li><li><p><strong>External service endpoints</strong>: Dev might point to sandbox APIs, production to live ones.</p></li></ul></blockquote><p>Things that should NOT differ between environments:</p><blockquote><ul><li><p><strong>Application code</strong>: The same Docker image runs everywhere. No environment-specific code paths.</p></li><li><p><strong>Business logic</strong>: If your app behaves differently in staging and production, you are going to have a bad time.</p></li><li><p><strong>Configuration structure</strong>: The same keys exist in all environments, just with different values.</p></li></ul></blockquote><p>Here is a practical structure using Parameter Store and Secrets Manager:</p><pre><code>Parameter Store:
  /dev/task-api/log-level       = "debug"
  /staging/task-api/log-level   = "info"
  /prod/task-api/log-level      = "warn"

  /dev/task-api/feature-new-ui  = "true"
  /staging/task-api/feature-new-ui = "true"
  /prod/task-api/feature-new-ui = "false"

Secrets Manager:
  dev/task-api/credentials      = { DB_URL: "...", API_KEY: "...", JWT_SECRET: "..." }
  staging/task-api/credentials  = { DB_URL: "...", API_KEY: "...", JWT_SECRET: "..." }
  prod/task-api/credentials     = { DB_URL: "...", API_KEY: "...", JWT_SECRET: "..." }</code></pre><p>Your ECS task definition sets a single environment variable, <code>APP_ENV</code>, to tell the application which environment it is running in. The config loader (like the one we built above) uses that value to fetch the right secrets:</p><pre><code>{
  "containerDefinitions": [
    {
      "name": "task-api",
      "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:v1.2.3",
      "environment": [
        { "name": "APP_ENV", "value": "prod" },
        { "name": "PORT", "value": "3000" }
      ]
    }
  ]
}</code></pre><p>Notice that the only values in the task definition are non-sensitive. The database URL and API keys are fetched from Secrets Manager at runtime, so they never appear in your Terraform code, CloudFormation templates, or ECS console.</p><h5><strong>ECS integration with Secrets Manager</strong></h5><p>ECS also has native integration with Secrets Manager, where it can inject secret values directly as environment variables when starting a container. This means your application does not need to call the Secrets Manager API at all:</p><pre><code>{
  "containerDefinitions": [
    {
      "name": "task-api",
      "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:v1.2.3",
      "secrets": [
        {
          "name": "DATABASE_URL",
          "valueFrom": "arn:aws:secretsmanager:us-east-1:123456789012:secret:prod/task-api/credentials:DB_URL::"
        },
        {
          "name": "API_KEY",
          "valueFrom": "arn:aws:secretsmanager:us-east-1:123456789012:secret:prod/task-api/credentials:API_KEY::"
        }
      ],
      "environment": [
        { "name": "APP_ENV", "value": "prod" },
        { "name": "PORT", "value": "3000" }
      ]
    }
  ]
}</code></pre><p>The <code>valueFrom</code> field uses the format <code>secret-arn:json-key:version-stage:version-id</code>. The double colon at the end means &#8220;use the latest version&#8221;. This approach is simpler because your application just reads <code>process.env.DATABASE_URL</code> like normal, and ECS handles the Secrets Manager integration.</p><p>The trade-off is that the secret values are only fetched when the container starts. If a secret rotates, you need to restart the container to pick up the new value. The SDK-based approach from the previous section lets you re-fetch secrets without restarting.</p><h5><strong>Advanced tools: Vault, SOPS, and Sealed Secrets</strong></h5><p>Everything we have covered so far handles the most common scenarios well. But as your infrastructure grows, you might need more specialized tools. I covered these in depth in the <a href="https://segfault.pw/blog/sre-secrets-management-in-kubernetes">SRE: Secrets Management in Kubernetes</a> article, so here is a quick overview with links:</p><blockquote><ul><li><p><strong><a href="https://www.vaultproject.io/">HashiCorp Vault</a></strong>: A full-featured secrets management platform. It supports dynamic secrets (generate a fresh database credential for each request), encryption as a service, and audit logging. Ideal for large organizations with complex compliance requirements.</p></li><li><p><strong><a href="https://github.com/getsops/sops">SOPS</a></strong>: Mozilla&#8217;s tool for encrypting secrets in files. You can store encrypted YAML, JSON, or .env files directly in Git. SOPS encrypts only the values, not the keys, so diffs are still readable. Great for GitOps workflows.</p></li><li><p><strong><a href="https://github.com/bitnami-labs/sealed-secrets">Sealed Secrets</a></strong>: A Kubernetes-specific solution. You encrypt secrets locally with a public key, commit the encrypted version to Git, and the Sealed Secrets controller in your cluster decrypts them. Perfect for GitOps with Kubernetes.</p></li></ul></blockquote><p>For the scope of this series, AWS Secrets Manager and Parameter Store will cover everything you need. If you are working with Kubernetes and want the deep dive into these tools, check out the SRE article linked above.</p><h5><strong>Quick reference: choosing the right approach</strong></h5><p>Here is a simple decision guide:</p><pre><code>Is it sensitive (password, API key, token)?
  YES --&gt; Use AWS Secrets Manager
    - Needs rotation? --&gt; Enable Secrets Manager rotation
    - Multiple services need it? --&gt; Use resource-based policy
  NO --&gt; Is it environment-specific config?
    YES --&gt; Use Parameter Store (free tier)
    NO --&gt; Hard-code it as an application default</code></pre><p>And here is a comparison table:</p><pre><code>Feature                  | Env Vars      | Parameter Store | Secrets Manager
-------------------------|---------------|-----------------|----------------
Cost                     | Free          | Free (std tier) | $0.40/secret/mo
Encryption               | No            | Optional (KMS)  | Always (KMS)
Rotation                 | Manual        | Manual          | Automatic
Audit logging            | No            | CloudTrail      | CloudTrail
Version history          | No            | Yes             | Yes
Cross-account access     | No            | Yes             | Yes
Best for                 | Local dev     | Non-sensitive    | Sensitive data</code></pre><h5><strong>Closing notes</strong></h5><p>You now have a solid understanding of how to manage configuration and secrets in a real application. The key takeaways are: follow the 12-factor methodology and keep config out of your code, use <code>.env</code> files for local development but never commit them, scan your repositories for leaked secrets with tools like gitleaks, use AWS Secrets Manager for sensitive values and Parameter Store for everything else, and structure your configuration so that the same Docker image works in every environment.</p><p>These are the fundamentals that will serve you well regardless of which cloud provider or orchestration platform you end up using. In the next article, we will tackle DNS, TLS, and making your application reachable from the internet with a proper domain name and HTTPS. See you there.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Deploying Your API to AWS ECS with Fargate]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-deploying-to-ecs</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-deploying-to-ecs</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!saW7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!saW7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!saW7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!saW7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!saW7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!saW7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!saW7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043139?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!saW7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!saW7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!saW7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!saW7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F63c00a98-369e-4e2b-b501-e5b1a7efaf8b_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article eight of the DevOps from Zero to Hero series. In the previous article we learned how to provision AWS infrastructure with Terraform. Now it is time to put that knowledge to work and deploy our TypeScript task API (the one we built in article two) to a real cloud environment using AWS ECS with Fargate.</p><p>ECS (Elastic Container Service) is AWS&#8217;s own container orchestration platform. It lets you run Docker containers without having to manage the underlying infrastructure yourself. When you pair it with Fargate, you do not even need to think about EC2 instances. You just define what your container needs, and AWS takes care of the rest. This is a great starting point before we get into Kubernetes later in the series.</p><p>In this article we will cover the core ECS concepts, push our Docker image to a registry, write Terraform code to provision everything (cluster, service, load balancer, auto-scaling), deploy the API, and verify it is running. By the end you will have a production-ready deployment that scales automatically based on demand.</p><p>Let&#8217;s get into it.</p><h5><strong>What is ECS?</strong></h5><p>Amazon Elastic Container Service (ECS) is a fully managed container orchestration service. Instead of installing and managing your own orchestrator (like Kubernetes), you hand your container definitions to ECS and it handles scheduling, scaling, and networking for you.</p><p>There are four key concepts you need to understand:</p><blockquote><ul><li><p><strong>Cluster</strong>: A logical grouping of resources where your containers run. Think of it as the boundary that holds everything together. A cluster can contain multiple services.</p></li><li><p><strong>Task Definition</strong>: A blueprint for your container. It specifies which Docker image to use, how much CPU and memory to allocate, what ports to expose, which environment variables to set, and where to send logs. It is versioned, so you can roll back to a previous definition if needed.</p></li><li><p><strong>Task</strong>: A running instance of a task definition. If the task definition is the recipe, the task is the actual dish being served. Each task runs one or more containers.</p></li><li><p><strong>Service</strong>: A long-running construct that ensures a specified number of tasks are always running. If a task crashes, the service automatically starts a new one. Services also handle rolling deployments when you update your task definition.</p></li></ul></blockquote><p>Here is how these pieces fit together:</p><pre><code>ECS Cluster
  &#9492;&#9472;&#9472; Service (maintains desired count of tasks)
        &#9500;&#9472;&#9472; Task 1 (running container based on task definition v3)
        &#9500;&#9472;&#9472; Task 2 (running container based on task definition v3)
        &#9492;&#9472;&#9472; Task 3 (running container based on task definition v3)</code></pre><h5><strong>ECS launch types: Fargate vs EC2</strong></h5><p>When you create an ECS service, you choose a launch type that determines where your containers actually run:</p><blockquote><ul><li><p><strong>EC2 launch type</strong>: You manage a fleet of EC2 instances. ECS schedules containers onto those instances. You are responsible for patching, scaling, and maintaining the instances. More control, more work.</p></li><li><p><strong>Fargate launch type</strong>: AWS manages the compute. You just specify CPU and memory for each task, and Fargate provisions the right amount of compute behind the scenes. No servers to manage, no capacity planning, no OS patches.</p></li></ul></blockquote><p>For this article we are using Fargate because it removes an entire layer of complexity. You pay a small premium compared to EC2, but you save a lot of operational effort. For most teams starting out, Fargate is the right choice.</p><h5><strong>ECS vs EKS: a brief comparison</strong></h5><p>You might wonder why we are not going straight to Kubernetes. AWS offers EKS (Elastic Kubernetes Service) for that. Here is the quick comparison:</p><blockquote><ul><li><p><strong>ECS</strong> is simpler to set up, tightly integrated with AWS services, and has no control plane cost with Fargate. If your workloads are AWS-only, ECS gets you running faster.</p></li><li><p><strong>EKS</strong> gives you the full Kubernetes API, portability across clouds, and access to the massive Kubernetes ecosystem. It is more complex but more flexible.</p></li></ul></blockquote><p>We will cover EKS in depth later in this series. For now, ECS with Fargate is the perfect stepping stone because it teaches you container orchestration concepts without the Kubernetes learning curve.</p><h5><strong>Pushing your Docker image to ECR</strong></h5><p>Before ECS can run your container, the image needs to be stored in a container registry that ECS can access. AWS provides ECR (Elastic Container Registry) for this purpose. You could also use GitHub Container Registry (GHCR) or Docker Hub, but ECR integrates seamlessly with ECS, so it is the simplest option.</p><p>First, create an ECR repository using the AWS CLI:</p><pre><code>aws ecr create-repository \
  --repository-name task-api \
  --region us-east-1 \
  --image-scanning-configuration scanOnPush=true</code></pre><p>The <code>scanOnPush=true</code> flag enables automatic vulnerability scanning on every push. This is a free feature and there is no reason not to use it.</p><p>Now authenticate Docker with ECR, build the image, tag it, and push:</p><pre><code># Get the login token and pipe it to docker login
aws ecr get-login-password --region us-east-1 | \
  docker login --username AWS --password-stdin \
  123456789012.dkr.ecr.us-east-1.amazonaws.com

# Build the image (using the Dockerfile from article 2)
docker build -t task-api .

# Tag it for ECR
docker tag task-api:latest \
  123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:latest

# Push to ECR
docker push \
  123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:latest</code></pre><p>Replace <code>123456789012</code> with your actual AWS account ID. You can find it by running <code>aws sts get-caller-identity --query Account --output text</code>.</p><h5><strong>The Terraform project structure</strong></h5><p>We are going to provision everything with Terraform, building on the foundations from article seven. Here is the project structure we will end up with:</p><pre><code>infra/
  &#9500;&#9472;&#9472; main.tf            # Provider and backend configuration
  &#9500;&#9472;&#9472; variables.tf       # Input variables
  &#9500;&#9472;&#9472; outputs.tf         # Output values
  &#9500;&#9472;&#9472; vpc.tf             # VPC, subnets, internet gateway
  &#9500;&#9472;&#9472; ecr.tf             # ECR repository
  &#9500;&#9472;&#9472; ecs.tf             # ECS cluster, task definition, service
  &#9500;&#9472;&#9472; alb.tf             # Application Load Balancer
  &#9500;&#9472;&#9472; autoscaling.tf     # Auto-scaling policies
  &#9500;&#9472;&#9472; iam.tf             # IAM roles and policies
  &#9492;&#9472;&#9472; security_groups.tf # Security groups</code></pre><p>Let&#8217;s start with the provider configuration and variables.</p><h5><strong>Provider and variables</strong></h5><p>The <code>main.tf</code> file configures the AWS provider and the Terraform backend:</p><pre><code># main.tf
terraform {
  required_version = "&gt;= 1.5.0"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~&gt; 5.0"
    }
  }

  backend "s3" {
    bucket = "my-terraform-state-bucket"
    key    = "ecs/task-api/terraform.tfstate"
    region = "us-east-1"
  }
}

provider "aws" {
  region = var.aws_region
}</code></pre><p>Now define the variables we will use throughout the configuration:</p><pre><code># variables.tf
variable "aws_region" {
  description = "AWS region to deploy to"
  type        = string
  default     = "us-east-1"
}

variable "project_name" {
  description = "Name of the project, used for resource naming"
  type        = string
  default     = "task-api"
}

variable "environment" {
  description = "Deployment environment"
  type        = string
  default     = "production"
}

variable "container_port" {
  description = "Port the container listens on"
  type        = number
  default     = 3000
}

variable "container_cpu" {
  description = "CPU units for the container (1024 = 1 vCPU)"
  type        = number
  default     = 256
}

variable "container_memory" {
  description = "Memory in MiB for the container"
  type        = number
  default     = 512
}

variable "desired_count" {
  description = "Number of tasks to run"
  type        = number
  default     = 2
}

variable "container_image" {
  description = "Docker image URI for the container"
  type        = string
}</code></pre><h5><strong>Networking: VPC and subnets</strong></h5><p>Our ECS service needs a VPC with public and private subnets. The ALB will sit in the public subnets, and the Fargate tasks will run in the private subnets:</p><pre><code># vpc.tf
data "aws_availability_zones" "available" {
  state = "available"
}

resource "aws_vpc" "main" {
  cidr_block           = "10.0.0.0/16"
  enable_dns_hostnames = true
  enable_dns_support   = true

  tags = {
    Name = "${var.project_name}-vpc"
  }
}

resource "aws_internet_gateway" "main" {
  vpc_id = aws_vpc.main.id

  tags = {
    Name = "${var.project_name}-igw"
  }
}

resource "aws_subnet" "public" {
  count                   = 2
  vpc_id                  = aws_vpc.main.id
  cidr_block              = "10.0.${count.index + 1}.0/24"
  availability_zone       = data.aws_availability_zones.available.names[count.index]
  map_public_ip_on_launch = true

  tags = {
    Name = "${var.project_name}-public-${count.index + 1}"
  }
}

resource "aws_subnet" "private" {
  count             = 2
  vpc_id            = aws_vpc.main.id
  cidr_block        = "10.0.${count.index + 10}.0/24"
  availability_zone = data.aws_availability_zones.available.names[count.index]

  tags = {
    Name = "${var.project_name}-private-${count.index + 1}"
  }
}

resource "aws_eip" "nat" {
  domain = "vpc"

  tags = {
    Name = "${var.project_name}-nat-eip"
  }
}

resource "aws_nat_gateway" "main" {
  allocation_id = aws_eip.nat.id
  subnet_id     = aws_subnet.public[0].id

  tags = {
    Name = "${var.project_name}-nat"
  }
}

resource "aws_route_table" "public" {
  vpc_id = aws_vpc.main.id

  route {
    cidr_block = "0.0.0.0/0"
    gateway_id = aws_internet_gateway.main.id
  }

  tags = {
    Name = "${var.project_name}-public-rt"
  }
}

resource "aws_route_table" "private" {
  vpc_id = aws_vpc.main.id

  route {
    cidr_block     = "0.0.0.0/0"
    nat_gateway_id = aws_nat_gateway.main.id
  }

  tags = {
    Name = "${var.project_name}-private-rt"
  }
}

resource "aws_route_table_association" "public" {
  count          = 2
  subnet_id      = aws_subnet.public[count.index].id
  route_table_id = aws_route_table.public.id
}

resource "aws_route_table_association" "private" {
  count          = 2
  subnet_id      = aws_subnet.private[count.index].id
  route_table_id = aws_route_table.private.id
}</code></pre><p>A few things to note here. The public subnets have a route to the internet gateway, which is where our ALB will live. The private subnets route through a NAT gateway, which lets our Fargate tasks pull images from ECR and send logs to CloudWatch without being directly exposed to the internet. This is a standard pattern for production workloads.</p><h5><strong>Security groups</strong></h5><p>We need two security groups: one for the ALB (allows inbound HTTP traffic from the internet) and one for the ECS tasks (allows traffic only from the ALB):</p><pre><code># security_groups.tf
resource "aws_security_group" "alb" {
  name        = "${var.project_name}-alb-sg"
  description = "Security group for the Application Load Balancer"
  vpc_id      = aws_vpc.main.id

  ingress {
    description = "HTTP from anywhere"
    from_port   = 80
    to_port     = 80
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }

  tags = {
    Name = "${var.project_name}-alb-sg"
  }
}

resource "aws_security_group" "ecs_tasks" {
  name        = "${var.project_name}-ecs-tasks-sg"
  description = "Security group for ECS tasks"
  vpc_id      = aws_vpc.main.id

  ingress {
    description     = "Allow traffic from ALB"
    from_port       = var.container_port
    to_port         = var.container_port
    protocol        = "tcp"
    security_groups = [aws_security_group.alb.id]
  }

  egress {
    from_port   = 0
    to_port     = 0
    protocol    = "-1"
    cidr_blocks = ["0.0.0.0/0"]
  }

  tags = {
    Name = "${var.project_name}-ecs-tasks-sg"
  }
}</code></pre><p>This is the principle of least privilege applied to networking. The ECS tasks only accept traffic from the ALB, not from the public internet directly. The ALB is the single entry point.</p><h5><strong>IAM roles for ECS</strong></h5><p>ECS tasks need two IAM roles: an execution role (used by ECS itself to pull images and write logs) and a task role (used by your application code to access AWS services):</p><pre><code># iam.tf
resource "aws_iam_role" "ecs_execution_role" {
  name = "${var.project_name}-ecs-execution-role"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Action = "sts:AssumeRole"
        Effect = "Allow"
        Principal = {
          Service = "ecs-tasks.amazonaws.com"
        }
      }
    ]
  })
}

resource "aws_iam_role_policy_attachment" "ecs_execution_role_policy" {
  role       = aws_iam_role.ecs_execution_role.name
  policy_arn = "arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy"
}

resource "aws_iam_role" "ecs_task_role" {
  name = "${var.project_name}-ecs-task-role"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Action = "sts:AssumeRole"
        Effect = "Allow"
        Principal = {
          Service = "ecs-tasks.amazonaws.com"
        }
      }
    ]
  })
}</code></pre><p>The execution role gets the managed <code>AmazonECSTaskExecutionRolePolicy</code>, which grants permissions to pull images from ECR and write logs to CloudWatch. The task role starts empty. As your application grows and needs access to other AWS services (S3, DynamoDB, SQS, etc.), you would attach policies to this role. Keep them separate so you maintain clear boundaries between what ECS needs and what your app needs.</p><h5><strong>The ECR repository in Terraform</strong></h5><p>Instead of creating the ECR repository manually with the CLI, let&#8217;s manage it with Terraform so everything is in code:</p><pre><code># ecr.tf
resource "aws_ecr_repository" "app" {
  name                 = var.project_name
  image_tag_mutability = "MUTABLE"
  force_delete         = true

  image_scanning_configuration {
    scan_on_push = true
  }

  tags = {
    Name = var.project_name
  }
}

resource "aws_ecr_lifecycle_policy" "app" {
  repository = aws_ecr_repository.app.name

  policy = jsonencode({
    rules = [
      {
        rulePriority = 1
        description  = "Keep only the last 10 images"
        selection = {
          tagStatus   = "any"
          countType   = "imageCountMoreThan"
          countNumber = 10
        }
        action = {
          type = "expire"
        }
      }
    ]
  })
}</code></pre><p>The lifecycle policy is important. Without it, your ECR repository will accumulate old images indefinitely, and you will pay for the storage. This policy keeps only the last 10 images and expires the rest automatically.</p><h5><strong>ECS cluster, task definition, and service</strong></h5><p>Now the main event. We are going to create the ECS cluster, define our task, and create a service that keeps it running:</p><pre><code># ecs.tf
resource "aws_cloudwatch_log_group" "app" {
  name              = "/ecs/${var.project_name}"
  retention_in_days = 30

  tags = {
    Name = var.project_name
  }
}

resource "aws_ecs_cluster" "main" {
  name = "${var.project_name}-cluster"

  setting {
    name  = "containerInsights"
    value = "enabled"
  }

  tags = {
    Name = "${var.project_name}-cluster"
  }
}

resource "aws_ecs_task_definition" "app" {
  family                   = var.project_name
  network_mode             = "awsvpc"
  requires_compatibilities = ["FARGATE"]
  cpu                      = var.container_cpu
  memory                   = var.container_memory
  execution_role_arn       = aws_iam_role.ecs_execution_role.arn
  task_role_arn            = aws_iam_role.ecs_task_role.arn

  container_definitions = jsonencode([
    {
      name      = var.project_name
      image     = var.container_image
      essential = true

      portMappings = [
        {
          containerPort = var.container_port
          protocol      = "tcp"
        }
      ]

      environment = [
        {
          name  = "NODE_ENV"
          value = "production"
        },
        {
          name  = "PORT"
          value = tostring(var.container_port)
        }
      ]

      logConfiguration = {
        logDriver = "awslogs"
        options = {
          "awslogs-group"         = aws_cloudwatch_log_group.app.name
          "awslogs-region"        = var.aws_region
          "awslogs-stream-prefix" = "ecs"
        }
      }

      healthCheck = {
        command     = ["CMD-SHELL", "curl -f http://localhost:${var.container_port}/health || exit 1"]
        interval    = 30
        timeout     = 5
        retries     = 3
        startPeriod = 60
      }
    }
  ])
}

resource "aws_ecs_service" "app" {
  name            = "${var.project_name}-service"
  cluster         = aws_ecs_cluster.main.id
  task_definition = aws_ecs_task_definition.app.arn
  desired_count   = var.desired_count
  launch_type     = "FARGATE"

  deployment_minimum_healthy_percent = 50
  deployment_maximum_percent         = 200
  health_check_grace_period_seconds  = 60

  network_configuration {
    subnets          = aws_subnet.private[*].id
    security_groups  = [aws_security_group.ecs_tasks.id]
    assign_public_ip = false
  }

  load_balancer {
    target_group_arn = aws_lb_target_group.app.arn
    container_name   = var.project_name
    container_port   = var.container_port
  }

  deployment_circuit_breaker {
    enable   = true
    rollback = true
  }

  depends_on = [aws_lb_listener.http]
}</code></pre><p>There is a lot happening here, so let&#8217;s break it down piece by piece.</p><p>The <strong>CloudWatch log group</strong> is where all container logs will be sent. Setting <code>retention_in_days</code> to 30 prevents logs from accumulating forever and running up your bill.</p><p>The <strong>cluster</strong> is straightforward. We enable Container Insights for better monitoring metrics.</p><p>The <strong>task definition</strong> is the most detailed part:</p><blockquote><ul><li><p><code>network_mode = "awsvpc"</code> gives each task its own elastic network interface. This is required for Fargate.</p></li><li><p><code>cpu</code> and <code>memory</code> define the Fargate sizing. 256 CPU units (0.25 vCPU) and 512 MiB is the smallest configuration and works well for a lightweight API.</p></li><li><p>The <code>container_definitions</code> block defines the container: image, port mappings, environment variables, log configuration, and health check.</p></li><li><p>The health check runs <code>curl</code> against the <code>/health</code> endpoint every 30 seconds. If three consecutive checks fail, ECS marks the task as unhealthy and replaces it.</p></li></ul></blockquote><p>The <strong>service</strong> ties everything together:</p><blockquote><ul><li><p><code>desired_count = 2</code> means ECS will always try to keep two tasks running.</p></li><li><p><code>deployment_minimum_healthy_percent = 50</code> means during a deployment, at least one task (50% of 2) must stay healthy. This allows rolling updates without downtime.</p></li><li><p><code>deployment_maximum_percent = 200</code> means ECS can temporarily run up to four tasks during a deployment (the old ones plus the new ones).</p></li><li><p>The <code>deployment_circuit_breaker</code> automatically rolls back a deployment if the new tasks fail to stabilize. This prevents a bad image from taking down your service.</p></li></ul></blockquote><h5><strong>Application Load Balancer</strong></h5><p>The ALB sits in front of your ECS service, distributes traffic across tasks, and provides a stable endpoint for clients. It also handles health checks to ensure traffic only goes to healthy tasks:</p><pre><code># alb.tf
resource "aws_lb" "app" {
  name               = "${var.project_name}-alb"
  internal           = false
  load_balancer_type = "application"
  security_groups    = [aws_security_group.alb.id]
  subnets            = aws_subnet.public[*].id

  tags = {
    Name = "${var.project_name}-alb"
  }
}

resource "aws_lb_target_group" "app" {
  name        = "${var.project_name}-tg"
  port        = var.container_port
  protocol    = "HTTP"
  vpc_id      = aws_vpc.main.id
  target_type = "ip"

  health_check {
    enabled             = true
    healthy_threshold   = 3
    unhealthy_threshold = 3
    timeout             = 5
    interval            = 30
    path                = "/health"
    protocol            = "HTTP"
    matcher             = "200"
  }

  deregistration_delay = 30

  tags = {
    Name = "${var.project_name}-tg"
  }
}

resource "aws_lb_listener" "http" {
  load_balancer_arn = aws_lb.app.arn
  port              = 80
  protocol          = "HTTP"

  default_action {
    type             = "forward"
    target_group_arn = aws_lb_target_group.app.arn
  }
}</code></pre><p>A few important details:</p><blockquote><ul><li><p><code>target_type = "ip"</code> is required for Fargate. With the EC2 launch type you would use <code>instance</code>, but Fargate tasks get their own IP addresses.</p></li><li><p>The health check hits <code>/health</code> and expects a <code>200</code> response. If a task fails three consecutive checks, the ALB stops sending it traffic and ECS replaces it.</p></li><li><p><code>deregistration_delay = 30</code> gives in-flight requests 30 seconds to complete before a task is removed from the target group during deployments. The default is 300 seconds, which is too long for most APIs.</p></li></ul></blockquote><p>In production you would add HTTPS support with an ACM certificate and a listener on port 443. We are keeping it simple with HTTP for now, but do not expose production APIs over plain HTTP.</p><h5><strong>Auto-scaling</strong></h5><p>Running a fixed number of tasks works, but it wastes money during low-traffic periods and risks overload during spikes. ECS integrates with Application Auto Scaling to adjust the task count based on metrics:</p><pre><code># autoscaling.tf
resource "aws_appautoscaling_target" "ecs" {
  max_capacity       = 10
  min_capacity       = 2
  resource_id        = "service/${aws_ecs_cluster.main.name}/${aws_ecs_service.app.name}"
  scalable_dimension = "ecs:service:DesiredCount"
  service_namespace  = "ecs"
}

resource "aws_appautoscaling_policy" "cpu" {
  name               = "${var.project_name}-cpu-scaling"
  policy_type        = "TargetTrackingScaling"
  resource_id        = aws_appautoscaling_target.ecs.resource_id
  scalable_dimension = aws_appautoscaling_target.ecs.scalable_dimension
  service_namespace  = aws_appautoscaling_target.ecs.service_namespace

  target_tracking_scaling_policy_configuration {
    predefined_metric_specification {
      predefined_metric_type = "ECSServiceAverageCPUUtilization"
    }
    target_value       = 70.0
    scale_in_cooldown  = 300
    scale_out_cooldown = 60
  }
}

resource "aws_appautoscaling_policy" "memory" {
  name               = "${var.project_name}-memory-scaling"
  policy_type        = "TargetTrackingScaling"
  resource_id        = aws_appautoscaling_target.ecs.resource_id
  scalable_dimension = aws_appautoscaling_target.ecs.scalable_dimension
  service_namespace  = aws_appautoscaling_target.ecs.service_namespace

  target_tracking_scaling_policy_configuration {
    predefined_metric_specification {
      predefined_metric_type = "ECSServiceAverageMemoryUtilization"
    }
    target_value       = 80.0
    scale_in_cooldown  = 300
    scale_out_cooldown = 60
  }
}</code></pre><p>Here is what this does:</p><blockquote><ul><li><p><strong>Minimum 2, maximum 10 tasks</strong>. You always have at least two tasks running for availability, and you cap at ten to control costs.</p></li><li><p><strong>CPU target: 70%</strong>. If average CPU across all tasks exceeds 70%, ECS adds more tasks. If it drops well below 70%, ECS removes tasks (down to the minimum of 2).</p></li><li><p><strong>Memory target: 80%</strong>. Same idea, but for memory utilization.</p></li><li><p><strong>Scale-out cooldown: 60 seconds</strong>. After adding tasks, wait at least 60 seconds before considering adding more. This prevents thrashing.</p></li><li><p><strong>Scale-in cooldown: 300 seconds</strong>. After removing tasks, wait 5 minutes before considering removing more. This is deliberately slower to avoid premature scale-down.</p></li></ul></blockquote><p>Target tracking is the simplest auto-scaling strategy and it works well for most workloads. You tell AWS &#8220;keep CPU around 70%&#8221; and it figures out how many tasks to run. If your scaling needs are more complex, you can use step scaling policies or scheduled scaling, but target tracking is a solid default.</p><h5><strong>Outputs</strong></h5><p>Finally, define outputs so you can easily find the ALB URL and other useful information after deploying:</p><pre><code># outputs.tf
output "alb_dns_name" {
  description = "DNS name of the Application Load Balancer"
  value       = aws_lb.app.dns_name
}

output "ecr_repository_url" {
  description = "URL of the ECR repository"
  value       = aws_ecr_repository.app.repository_url
}

output "ecs_cluster_name" {
  description = "Name of the ECS cluster"
  value       = aws_ecs_cluster.main.name
}

output "ecs_service_name" {
  description = "Name of the ECS service"
  value       = aws_ecs_service.app.name
}

output "cloudwatch_log_group" {
  description = "CloudWatch log group for the ECS tasks"
  value       = aws_cloudwatch_log_group.app.name
}</code></pre><h5><strong>Deploying with Terraform</strong></h5><p>With all the configuration in place, deploying is a matter of running the standard Terraform workflow:</p><pre><code>cd infra

# Initialize Terraform (download providers, configure backend)
terraform init

# Review the execution plan
terraform plan -var="container_image=123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:latest"

# Apply the changes
terraform apply -var="container_image=123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:latest"</code></pre><p>Terraform will show you everything it plans to create before it does anything. Review the plan carefully, then type <code>yes</code> to proceed. The first deployment takes a few minutes because it needs to create the VPC, subnets, NAT gateway, ALB, and ECS resources.</p><p>When it finishes, Terraform will print the outputs. Grab the <code>alb_dns_name</code> value, that is your API endpoint.</p><h5><strong>Testing the deployment</strong></h5><p>Let&#8217;s verify everything is working. Use the ALB DNS name from the Terraform output:</p><pre><code># Check the health endpoint
curl http://task-api-alb-123456789.us-east-1.elb.amazonaws.com/health

# Expected response:
# {"status":"healthy","uptime":42.123,"timestamp":"2026-05-12T10:30:00.000Z"}</code></pre><p>Try creating a task:</p><pre><code># Create a new task
curl -X POST \
  http://task-api-alb-123456789.us-east-1.elb.amazonaws.com/tasks \
  -H "Content-Type: application/json" \
  -d '{"title": "Deploy to ECS", "description": "First task from production!"}'

# List all tasks
curl http://task-api-alb-123456789.us-east-1.elb.amazonaws.com/tasks</code></pre><p>If everything is working, you should see healthy responses. If something is wrong, check the CloudWatch logs:</p><pre><code># View recent logs from the ECS tasks
aws logs tail /ecs/task-api --follow --since 10m</code></pre><p>You can also check the ECS service events to see if tasks are starting and stopping properly:</p><pre><code>aws ecs describe-services \
  --cluster task-api-cluster \
  --services task-api-service \
  --query 'services[0].events[:10]' \
  --output table</code></pre><h5><strong>Deploying updates: the rolling deployment flow</strong></h5><p>When you push a new version of your Docker image, you need to tell ECS to pick it up. The simplest way is to force a new deployment:</p><pre><code># Build, tag, and push the new image
docker build -t task-api .
docker tag task-api:latest \
  123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:latest
docker push \
  123456789012.dkr.ecr.us-east-1.amazonaws.com/task-api:latest

# Force ECS to pull the new image
aws ecs update-service \
  --cluster task-api-cluster \
  --service task-api-service \
  --force-new-deployment</code></pre><p>Here is what happens during a rolling deployment:</p><blockquote><ul><li><p>ECS starts new tasks with the updated image alongside the existing ones (up to <code>deployment_maximum_percent</code>)</p></li><li><p>The ALB health checks verify the new tasks are healthy</p></li><li><p>Once the new tasks pass health checks, the ALB starts routing traffic to them</p></li><li><p>ECS drains connections from the old tasks (respecting <code>deregistration_delay</code>)</p></li><li><p>The old tasks are stopped</p></li><li><p>If the new tasks fail to become healthy, the deployment circuit breaker automatically rolls back to the previous version</p></li></ul></blockquote><p>This entire process happens with zero downtime. Your users never see an error during the deployment because the old tasks keep serving traffic until the new ones are ready.</p><p>In a real CI/CD pipeline (which we covered earlier in the series), you would automate this entire flow. Push to main, CI builds the image, pushes to ECR, and triggers the ECS deployment. No manual steps required.</p><h5><strong>Cost considerations</strong></h5><p>Before we wrap up, let&#8217;s talk about what this costs. Fargate pricing is based on the CPU and memory you allocate to each task, billed per second with a one-minute minimum:</p><blockquote><ul><li><p><strong>0.25 vCPU, 512 MiB</strong> (our configuration): roughly $0.01/hour per task</p></li><li><p><strong>With 2 tasks running 24/7</strong>: approximately $15/month for compute</p></li><li><p><strong>NAT gateway</strong>: about $32/month (this is often the largest cost for small deployments)</p></li><li><p><strong>ALB</strong>: approximately $16/month plus data transfer</p></li><li><p><strong>ECR</strong>: $0.10/GB/month for storage, first 500 MB free</p></li><li><p><strong>CloudWatch Logs</strong>: $0.50/GB ingested</p></li></ul></blockquote><p>For a small API, you are looking at roughly $65-80/month total. The NAT gateway is the single most expensive component. If cost is a concern, you could run your tasks in public subnets with <code>assign_public_ip = true</code> and skip the NAT gateway, but this is not recommended for production workloads because it exposes your tasks directly to the internet.</p><h5><strong>Closing notes</strong></h5><p>You now have a production-grade deployment of your TypeScript API on AWS ECS with Fargate. The setup includes a proper VPC with public and private subnets, an Application Load Balancer for traffic distribution and health checking, auto-scaling to handle variable load, a deployment circuit breaker for safety, and centralized logging in CloudWatch. All managed as code with Terraform.</p><p>ECS with Fargate is a great choice when you want container orchestration without the complexity of Kubernetes. It integrates tightly with the AWS ecosystem, requires minimal operational overhead, and scales well for most workloads.</p><p>In the next article, we will look at more advanced AWS services and prepare for the jump to Kubernetes with EKS. If you have followed along this far, you already understand the fundamentals of container orchestration, which will make Kubernetes much easier to learn.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Infrastructure as Code with Terraform]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-infrastructure-as-code</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-infrastructure-as-code</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Sat, 09 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rB-Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rB-Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rB-Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!rB-Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!rB-Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!rB-Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rB-Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/efbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043140?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rB-Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!rB-Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!rB-Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!rB-Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fefbe4bbc-0e40-4949-a8f9-db22992fb19a_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article seven of the DevOps from Zero to Hero series. In the previous article we explored AWS networking: VPCs, subnets, route tables, and security groups. Now it is time to stop clicking around in the AWS console and start defining infrastructure the same way we define application code: in files, under version control, with repeatable results.</p><p>This is Infrastructure as Code (IaC), and it is one of the most important practices in modern DevOps. If you have ever manually created an EC2 instance, realized you forgot a tag, created another one differently, and then had no idea which was &#8220;the right one,&#8221; you already understand the problem IaC solves.</p><p>We will cover what IaC is, walk through the core Terraform workflow, learn how to manage state safely, and build a real VPC with public and private subnets using HCL files. If you want to go deeper after this, check out <a href="https://segfault.pw/blog/getting_started_with_terraform_modules">Getting started with Terraform modules</a> and <a href="https://segfault.pw/blog/brief_introduction_to_terratest">Brief introduction to Terratest</a>.</p><p>Let&#8217;s get into it.</p><h5><strong>What is Infrastructure as Code?</strong></h5><p>IaC means defining your infrastructure (servers, networks, databases, load balancers, DNS records) in declarative configuration files rather than creating them manually through a web console.</p><blockquote><ul><li><p><strong>Reproducibility</strong>: Recreate your entire infrastructure from scratch with a single command. No more &#8220;it works in staging but not in production&#8221; because someone configured something differently.</p></li><li><p><strong>Version control</strong>: Every change is tracked in Git. You can see who changed what, when, and why.</p></li><li><p><strong>Collaboration</strong>: Infrastructure changes go through pull requests just like code changes.</p></li><li><p><strong>Drift detection</strong>: IaC tools detect when real state drifts from declared state and bring it back in line.</p></li><li><p><strong>Documentation</strong>: Your code IS your documentation. Always up to date because it is the source of truth.</p></li></ul></blockquote><h5><strong>IaC vs ClickOps</strong></h5><p>&#8220;ClickOps&#8221; is the term for managing infrastructure by clicking through a cloud console. It is fine for learning but falls apart in teams:</p><blockquote><ul><li><p><strong>No audit trail</strong>: Someone changes a security group rule. Three months later, nobody remembers who or why.</p></li><li><p><strong>Snowflake servers</strong>: Each environment is slightly different because different people configured them at different times.</p></li><li><p><strong>No reproducibility</strong>: Could you recreate your production environment from scratch? How long would it take?</p></li><li><p><strong>Human error</strong>: At 2 AM you accidentally delete a production database because you were in the wrong tab.</p></li><li><p><strong>Knowledge silos</strong>: Only one person knows how the network is configured because they set it up manually.</p></li></ul></blockquote><p>IaC eliminates all of these problems. Infrastructure defined in code, reviewed by the team, tracked in Git, reproducible at any time.</p><h5><strong>Why Terraform?</strong></h5><p>Several IaC tools exist:</p><blockquote><ul><li><p><strong>CloudFormation</strong>: AWS-native, JSON/YAML. AWS-only, verbose, but deep AWS integration.</p></li><li><p><strong>Pulumi</strong>: Infrastructure in real programming languages (TypeScript, Python, Go). Great DX, smaller community.</p></li><li><p><strong>AWS CDK</strong>: Generates CloudFormation using TypeScript or Python. AWS-only, nicer than raw CloudFormation.</p></li><li><p><strong>Terraform</strong>: HashiCorp&#8217;s tool using HCL. Works across AWS, GCP, Azure, Kubernetes, and hundreds of providers.</p></li></ul></blockquote><p>We use Terraform because it works across clouds, has the largest ecosystem, and is what most teams use. The concepts (state, plans, declarative config) transfer to any IaC tool.</p><h5><strong>Terraform basics: the building blocks</strong></h5><p>Terraform uses HCL (HashiCorp Configuration Language), a declarative language for describing infrastructure.</p><p><strong>Providers</strong> are plugins that let Terraform talk to a cloud or service:</p><pre><code>terraform {
  required_version = "&gt;= 1.0"
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~&gt; 5.0"
    }
  }
}

provider "aws" {
  region = "us-east-1"
}</code></pre><p><strong>Resources</strong> describe a piece of infrastructure:</p><pre><code>resource "aws_instance" "web" {
  ami           = "ami-0c55b159cbfafe1f0"
  instance_type = "t3.micro"
  tags = { Name = "web-server" }
}</code></pre><p><strong>Data sources</strong> read information without creating anything:</p><pre><code>data "aws_ami" "ubuntu" {
  most_recent = true
  filter {
    name   = "name"
    values = ["ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-*"]
  }
  owners = ["099720109477"]
}</code></pre><p><strong>Variables</strong> parameterize your configuration:</p><pre><code>variable "environment" {
  description = "Deployment environment"
  type        = string
  default     = "dev"
}

variable "instance_type" {
  description = "EC2 instance type"
  type        = string
  default     = "t3.micro"
}</code></pre><p><strong>Outputs</strong> extract values after creation:</p><pre><code>output "instance_public_ip" {
  description = "Public IP of the web server"
  value       = aws_instance.web.public_ip
}</code></pre><h5><strong>The Terraform workflow: init, plan, apply, destroy</strong></h5><p><strong>terraform init</strong> downloads providers and sets up the backend:</p><pre><code>$ terraform init
Initializing provider plugins...
- Installing hashicorp/aws v5.82.1...
Terraform has been successfully initialized!</code></pre><p><strong>terraform plan</strong> shows what would change without changing anything:</p><pre><code>$ terraform plan
  # aws_instance.web will be created
  + resource "aws_instance" "web" {
      + ami           = "ami-0c55b159cbfafe1f0"
      + instance_type = "t3.micro"
    }
Plan: 1 to add, 0 to change, 0 to destroy.</code></pre><p>The symbols: <code>+</code> create, <code>~</code> modify, <code>-</code> destroy, <code>-/+</code> replace. Always read the plan before applying.</p><p><strong>terraform apply</strong> makes changes real (asks for confirmation):</p><pre><code>$ terraform apply
aws_instance.web: Creating...
aws_instance.web: Creation complete after 32s [id=i-0abc123def456789]
Apply complete! Resources: 1 added, 0 changed, 0 destroyed.</code></pre><p><strong>terraform destroy</strong> tears everything down when you no longer need it.</p><h5><strong>State management</strong></h5><p>Terraform records what it created in a state file. By default this is local (<code>terraform.tfstate</code>), which breaks in teams:</p><blockquote><ul><li><p><strong>No sharing</strong>: Teammates cannot run Terraform without the state file.</p></li><li><p><strong>No locking</strong>: Two concurrent applies can corrupt state or create duplicates.</p></li><li><p><strong>Risk of loss</strong>: Laptop dies, state is gone, Terraform forgets your infrastructure.</p></li></ul></blockquote><p>The solution is remote state with S3 + DynamoDB locking:</p><pre><code># Create S3 bucket for state (one-time setup)
aws s3api create-bucket --bucket my-terraform-state --region us-east-1
aws s3api put-bucket-versioning --bucket my-terraform-state \
  --versioning-configuration Status=Enabled

# Create DynamoDB table for locking
aws dynamodb create-table --table-name terraform-lock \
  --attribute-definitions AttributeName=LockID,AttributeType=S \
  --key-schema AttributeName=LockID,KeyType=HASH \
  --billing-mode PAY_PER_REQUEST --region us-east-1</code></pre><p>Then configure the backend:</p><pre><code>terraform {
  backend "s3" {
    bucket         = "my-terraform-state"
    key            = "prod/network/terraform.tfstate"
    region         = "us-east-1"
    dynamodb_table = "terraform-lock"
    encrypt        = true
  }
}</code></pre><p>Now state is shared, versioned, encrypted, and locked during applies.</p><h5><strong>Practical example: provisioning a VPC</strong></h5><p>Let&#8217;s build a VPC with public and private subnets, an internet gateway, route tables, and a security group, the same architecture from the networking article, but as code.</p><p><strong>variables.tf</strong></p><pre><code>variable "aws_region" {
  description = "AWS region"
  type        = string
  default     = "us-east-1"
}

variable "environment" {
  description = "Environment name"
  type        = string
  default     = "dev"
}

variable "vpc_cidr" {
  description = "CIDR block for the VPC"
  type        = string
  default     = "10.0.0.0/16"
}

variable "public_subnet_cidrs" {
  type    = list(string)
  default = ["10.0.1.0/24", "10.0.2.0/24"]
}

variable "private_subnet_cidrs" {
  type    = list(string)
  default = ["10.0.10.0/24", "10.0.11.0/24"]
}

variable "allowed_ssh_cidr" {
  type    = string
  default = "0.0.0.0/0"
}</code></pre><p><strong>main.tf</strong></p><pre><code>data "aws_availability_zones" "available" {
  state = "available"
}

resource "aws_vpc" "main" {
  cidr_block           = var.vpc_cidr
  enable_dns_support   = true
  enable_dns_hostnames = true
  tags = {
    Name = "${var.environment}-vpc"
    Environment = var.environment
    ManagedBy   = "terraform"
  }
}

resource "aws_internet_gateway" "main" {
  vpc_id = aws_vpc.main.id
  tags   = { Name = "${var.environment}-igw" }
}

resource "aws_subnet" "public" {
  count                   = length(var.public_subnet_cidrs)
  vpc_id                  = aws_vpc.main.id
  cidr_block              = var.public_subnet_cidrs[count.index]
  availability_zone       = data.aws_availability_zones.available.names[count.index]
  map_public_ip_on_launch = true
  tags = { Name = "${var.environment}-public-${count.index + 1}", Tier = "public" }
}

resource "aws_subnet" "private" {
  count             = length(var.private_subnet_cidrs)
  vpc_id            = aws_vpc.main.id
  cidr_block        = var.private_subnet_cidrs[count.index]
  availability_zone = data.aws_availability_zones.available.names[count.index]
  tags = { Name = "${var.environment}-private-${count.index + 1}", Tier = "private" }
}

resource "aws_route_table" "public" {
  vpc_id = aws_vpc.main.id
  route {
    cidr_block = "0.0.0.0/0"
    gateway_id = aws_internet_gateway.main.id
  }
  tags = { Name = "${var.environment}-public-rt" }
}

resource "aws_route_table_association" "public" {
  count          = length(var.public_subnet_cidrs)
  subnet_id      = aws_subnet.public[count.index].id
  route_table_id = aws_route_table.public.id
}

resource "aws_route_table" "private" {
  vpc_id = aws_vpc.main.id
  tags   = { Name = "${var.environment}-private-rt" }
}

resource "aws_route_table_association" "private" {
  count          = length(var.private_subnet_cidrs)
  subnet_id      = aws_subnet.private[count.index].id
  route_table_id = aws_route_table.private.id
}

resource "aws_security_group" "web" {
  name        = "${var.environment}-web-sg"
  description = "Allow HTTP, HTTPS, and SSH"
  vpc_id      = aws_vpc.main.id
  tags = { Name = "${var.environment}-web-sg" }
}

resource "aws_vpc_security_group_ingress_rule" "http" {
  security_group_id = aws_security_group.web.id
  cidr_ipv4 = "0.0.0.0/0"
  from_port = 80
  to_port   = 80
  ip_protocol = "tcp"
}

resource "aws_vpc_security_group_ingress_rule" "https" {
  security_group_id = aws_security_group.web.id
  cidr_ipv4 = "0.0.0.0/0"
  from_port = 443
  to_port   = 443
  ip_protocol = "tcp"
}

resource "aws_vpc_security_group_ingress_rule" "ssh" {
  security_group_id = aws_security_group.web.id
  cidr_ipv4 = var.allowed_ssh_cidr
  from_port = 22
  to_port   = 22
  ip_protocol = "tcp"
}

resource "aws_vpc_security_group_egress_rule" "all_outbound" {
  security_group_id = aws_security_group.web.id
  cidr_ipv4   = "0.0.0.0/0"
  ip_protocol = "-1"
}</code></pre><p><strong>outputs.tf</strong></p><pre><code>output "vpc_id" {
  value = aws_vpc.main.id
}

output "public_subnet_ids" {
  value = aws_subnet.public[*].id
}

output "private_subnet_ids" {
  value = aws_subnet.private[*].id
}

output "security_group_id" {
  value = aws_security_group.web.id
}</code></pre><p>What this creates:</p><blockquote><ul><li><p><strong>VPC</strong> with DNS support using the configured CIDR block</p></li><li><p><strong>Internet Gateway</strong> attached to the VPC for public internet access</p></li><li><p><strong>Public subnets</strong> across availability zones with automatic public IP assignment</p></li><li><p><strong>Private subnets</strong> with no internet route, keeping resources isolated</p></li><li><p><strong>Route tables</strong> directing public traffic through the gateway</p></li><li><p><strong>Security group</strong> allowing HTTP, HTTPS, SSH inbound and all outbound</p></li></ul></blockquote><h5><strong>Variables and tfvars</strong></h5><p>Use <code>.tfvars</code> files for environment-specific values:</p><pre><code># terraform.tfvars (dev defaults)
aws_region  = "us-east-1"
environment = "dev"
vpc_cidr    = "10.0.0.0/16"

# prod.tfvars
# aws_region           = "us-east-1"
# environment          = "prod"
# vpc_cidr             = "10.1.0.0/16"
# public_subnet_cidrs  = ["10.1.1.0/24", "10.1.2.0/24", "10.1.3.0/24"]
# private_subnet_cidrs = ["10.1.10.0/24", "10.1.11.0/24", "10.1.12.0/24"]
# allowed_ssh_cidr     = "203.0.113.0/24"</code></pre><pre><code># Uses terraform.tfvars automatically
terraform plan

# Uses a specific file
terraform plan -var-file="prod.tfvars"

# Or pass directly
terraform plan -var="environment=staging"

# Or use environment variables
export TF_VAR_environment="staging"</code></pre><p>Precedence (lowest to highest): defaults, <code>terraform.tfvars</code>, <code>*.auto.tfvars</code>, <code>-var-file</code>, <code>-var</code>, <code>TF_VAR_</code> env vars.</p><h5><strong>Running the example</strong></h5><pre><code>terraform init       # Download providers
terraform fmt        # Format code
terraform validate   # Check syntax
terraform plan       # Preview changes
terraform apply      # Create infrastructure
terraform output     # Show outputs
terraform state list # List managed resources
terraform destroy    # Clean up when done</code></pre><h5><strong>Best practices</strong></h5><p>A few things to keep in mind:</p><blockquote><ul><li><p><strong>Never commit state files</strong> to Git. They contain sensitive data. Use remote state.</p></li><li><p><strong>Do commit <code>.terraform.lock.hcl</code></strong>. It pins provider versions like <code>package-lock.json</code>.</p></li><li><p><strong>Be careful with <code>.tfvars</code></strong>. If they contain secrets, use environment variables or a secrets manager instead.</p></li><li><p><strong>Tag everything</strong> with <code>ManagedBy = "terraform"</code> so you can distinguish IaC-managed resources from manual ones.</p></li><li><p><strong>Use <code>plan -out=tfplan</code></strong> in CI/CD to save a plan file and apply exactly what was reviewed.</p></li></ul></blockquote><h5><strong>Closing notes</strong></h5><p>Infrastructure as Code changes how you think about infrastructure. Instead of fragile, manually configured environments, you get reproducible, version-controlled definitions that anyone on the team can read and modify.</p><p>Terraform is not the only tool, but it is a great starting point. The declarative approach (describe what you want, Terraform figures out how to get there) makes it accessible, and the plan-before-apply workflow gives you a safety net that clicking through a console never could.</p><p>Start small. One resource, one plan, one apply. Then add more. Before long your entire infrastructure lives in a handful of files and you will wonder how you ever managed without it.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: AWS from Scratch]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-aws-from-scratch</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-aws-from-scratch</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Wed, 06 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!MLuO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MLuO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MLuO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!MLuO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!MLuO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!MLuO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MLuO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MLuO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!MLuO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!MLuO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!MLuO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbb35eca3-5f4a-4b25-a91b-008742447a2f_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article six of the DevOps from Zero to Hero series. We have covered DevOps concepts, built a TypeScript API, learned version control, automated testing, and set up CI/CD. Now it is time to talk about the cloud. We will set up an AWS account the right way, configure IAM with least privilege, install the AWS CLI, understand VPC networking and security groups, and overview the key services we will use throughout the rest of this series.</p><p>Let&#8217;s get into it.</p><h5><strong>Creating your AWS account</strong></h5><p>Go to <a href="https://aws.amazon.com">aws.amazon.com</a> and click &#8220;Create an AWS Account.&#8221; You need an email, a credit card, and a phone number. AWS will not charge you for creating the account, and many services have a free tier for 12 months. When you create the account, you get a root user with unrestricted access to everything. Never use it for daily work. Here is what to do immediately:</p><blockquote><ul><li><p><strong>Enable MFA on root</strong>: Go to IAM, click on the root user, and set up multi-factor authentication with an authenticator app like Google Authenticator or Authy</p></li><li><p><strong>Create an IAM admin user</strong>: We will cover this next. From here on, you log in with the IAM user, not root</p></li><li><p><strong>Set up a billing alarm</strong>: Go to CloudWatch and create an alarm when estimated charges exceed $10</p></li><li><p><strong>Store root credentials securely</strong>: Save the password and MFA recovery codes in a password manager, then stop using root</p></li></ul></blockquote><h5><strong>IAM: Identity and Access Management</strong></h5><p>IAM controls who can do what in your AWS account. There are four main concepts:</p><blockquote><ul><li><p><strong>Users</strong>: Individual people or applications, each with their own credentials</p></li><li><p><strong>Groups</strong>: Collections of users. Attach permissions to the group, then add users to it</p></li><li><p><strong>Roles</strong>: Temporary identities assumed by users, services, or applications (e.g., an EC2 instance assuming a role to read from S3)</p></li><li><p><strong>Policies</strong>: JSON documents defining what actions are allowed or denied on which resources</p></li></ul></blockquote><p><strong>The principle of least privilege</strong>: give every identity only the permissions it needs. If credentials get leaked, the blast radius is limited to what that identity was allowed to do.</p><p><strong>Creating an IAM admin user:</strong></p><ol><li><p>Go to IAM console, click &#8220;Users,&#8221; then &#8220;Create user.&#8221;</p></li><li><p>Name it <code>admin</code>, check &#8220;Provide user access to the AWS Management Console.&#8221;</p></li><li><p>Attach the <code>AdministratorAccess</code> policy directly.</p></li><li><p>Save the sign-in URL and credentials, log in as this user, and enable MFA.</p></li></ol><p><strong>Writing a custom IAM policy:</strong></p><p>Here is a least-privilege policy allowing upload and download from a specific S3 bucket:</p><pre><code>{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowS3ReadWrite",
      "Effect": "Allow",
      "Action": [
        "s3:GetObject",
        "s3:PutObject",
        "s3:ListBucket"
      ],
      "Resource": [
        "arn:aws:s3:::my-app-uploads",
        "arn:aws:s3:::my-app-uploads/*"
      ]
    }
  ]
}</code></pre><blockquote><ul><li><p><strong>Version</strong>: Always <code>"2012-10-17"</code> (the policy language version, not a date you change)</p></li><li><p><strong>Effect</strong>: <code>"Allow"</code> or <code>"Deny"</code>. Deny always wins</p></li><li><p><strong>Action</strong>: Specific API calls being permitted</p></li><li><p><strong>Resource</strong>: AWS resources identified by ARN. Note we need both the bucket and <code>/*</code> for objects inside it</p></li></ul></blockquote><p><strong>IAM best practices:</strong></p><blockquote><ul><li><p><strong>Never use root</strong> except for billing and account emergencies</p></li><li><p><strong>Use groups</strong> to manage permissions, not individual user policies</p></li><li><p><strong>Use roles instead of access keys</strong> for EC2, Lambda, and ECS</p></li><li><p><strong>Enable MFA</strong> on every human user</p></li><li><p><strong>Rotate access keys</strong> every 90 days</p></li></ul></blockquote><h5><strong>AWS CLI setup and configuration</strong></h5><p>The AWS CLI lets you interact with AWS from your terminal. Install it:</p><pre><code># Linux
curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o "awscliv2.zip"
unzip awscliv2.zip
sudo ./aws/install
aws --version

# macOS
curl "https://awscli.amazonaws.com/AWSCLIV2.pkg" -o "AWSCLIV2.pkg"
sudo installer -pkg AWSCLIV2.pkg -target /
aws --version</code></pre><p>Create an access key in IAM for your admin user, then configure:</p><pre><code>aws configure
# AWS Access Key ID [None]: AKIAIOSFODNN7EXAMPLE
# AWS Secret Access Key [None]: wJalrXUtnFEMI/K7MDENG/bPxRfiCYEXAMPLEKEY
# Default region name [None]: us-east-1
# Default output format [None]: json</code></pre><p><strong>Named profiles</strong> let you manage multiple accounts (dev, staging, prod):</p><pre><code>aws configure --profile dev
aws configure --profile staging
aws configure --profile prod

# Use a specific profile
aws s3 ls --profile dev

# Or set via environment variable
export AWS_PROFILE=dev
aws sts get-caller-identity</code></pre><p>Set your default to dev so you never accidentally run commands against production.</p><h5><strong>VPC: Virtual Private Cloud</strong></h5><p>A VPC is your own isolated network inside AWS. Every resource you launch lives inside a VPC. You get a default VPC in each region, but for production create custom VPCs with explicit control.</p><p><strong>Key concepts:</strong></p><blockquote><ul><li><p><strong>CIDR block</strong>: The IP range for your VPC (e.g., <code>10.0.0.0/16</code> gives 65,536 IPs). Chosen at creation, cannot be changed later</p></li><li><p><strong>Subnets</strong>: Subdivisions of your VPC, each in a specific availability zone</p></li><li><p><strong>Availability zones (AZs)</strong>: Physically separate data centers within a region for fault tolerance</p></li></ul></blockquote><p><strong>Public vs private subnets:</strong></p><blockquote><ul><li><p><strong>Public subnets</strong> have a route to the internet through an Internet Gateway. Use for load balancers and bastion hosts</p></li><li><p><strong>Private subnets</strong> have no direct internet route. Use for app servers and databases. They reach the internet outbound through a NAT gateway</p></li></ul></blockquote><p><strong>Route tables</strong> tell traffic where to go. Public subnets route <code>0.0.0.0/0</code> to the Internet Gateway; private subnets route it through the NAT Gateway. Here is a typical production VPC:</p><pre><code>                        Region: us-east-1
 &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;
 &#9474;                    VPC: 10.0.0.0/16                      &#9474;
 &#9474;                                                          &#9474;
 &#9474;  &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;    &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488;      &#9474;
 &#9474;  &#9474;   AZ: us-east-1a    &#9474;    &#9474;   AZ: us-east-1b    &#9474;      &#9474;
 &#9474;  &#9474;                     &#9474;    &#9474;                     &#9474;      &#9474;
 &#9474;  &#9474; &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488; &#9474;    &#9474; &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; Public Subnet   &#9474; &#9474;    &#9474; &#9474; Public Subnet   &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; 10.0.1.0/24     &#9474; &#9474;    &#9474; &#9474; 10.0.2.0/24     &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; [Load Balancer] &#9474; &#9474;    &#9474; &#9474; [NAT Gateway]   &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496; &#9474;    &#9474; &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496; &#9474;      &#9474;
 &#9474;  &#9474; &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488; &#9474;    &#9474; &#9484;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9488; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; Private Subnet  &#9474; &#9474;    &#9474; &#9474; Private Subnet  &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; 10.0.3.0/24     &#9474; &#9474;    &#9474; &#9474; 10.0.4.0/24     &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; [App Server]    &#9474; &#9474;    &#9474; &#9474; [App Server]    &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9474; [Database]      &#9474; &#9474;    &#9474; &#9474; [Database]      &#9474; &#9474;      &#9474;
 &#9474;  &#9474; &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496; &#9474;    &#9474; &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496; &#9474;      &#9474;
 &#9474;  &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;    &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;      &#9474;
 &#9474;                                                          &#9474;
 &#9474;                    [Internet Gateway]                     &#9474;
 &#9492;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9496;
                            &#9474;
                        Internet</code></pre><p>High availability (two AZs), security (databases in private subnets), and outbound access through NAT. We will build this VPC with Terraform in a later article.</p><h5><strong>Security groups</strong></h5><p>Security groups are virtual firewalls for your resources. They are stateful (allow inbound on port 80 and the response is automatically allowed outbound), allow-only (no deny rules, unlisted traffic is denied), and instance-level (attached to resources, not subnets).</p><p><strong>Web server security group example:</strong></p><pre><code>Security Group: web-server-sg

Inbound Rules:
  HTTP       TCP  80     0.0.0.0/0       Allow HTTP from anywhere
  HTTPS      TCP  443    0.0.0.0/0       Allow HTTPS from anywhere
  SSH        TCP  22     203.0.113.50/32 Allow SSH from my IP only

Outbound Rules:
  All        All  All    0.0.0.0/0       Allow all outbound</code></pre><p><strong>Database security group referencing the web server group:</strong></p><pre><code>Security Group: database-sg

Inbound Rules:
  PostgreSQL TCP  5432   web-server-sg   Allow Postgres from web servers only

Outbound Rules:
  All        All  All    0.0.0.0/0       Allow all outbound</code></pre><p>The database references <code>web-server-sg</code> as the source, so any instance with that group can reach the database. No IP tracking needed. Best practices: never open SSH to <code>0.0.0.0/0</code>, use security group references instead of IPs for internal communication, and create separate groups per role.</p><h5><strong>Key AWS services overview</strong></h5><p>AWS has over 200 services. Here are the core ones for this series:</p><blockquote><ul><li><p><strong>EC2</strong>: Virtual servers. Choose OS, CPU, memory. Pay by the second</p></li><li><p><strong>S3</strong>: Object storage for files. Used for assets, backups, logs, static hosting. 99.999999999% durability</p></li><li><p><strong>RDS</strong>: Managed databases (PostgreSQL, MySQL, etc.). AWS handles backups, patching, failover</p></li><li><p><strong>ECS</strong>: Runs Docker containers on EC2 or Fargate (serverless). Handles scheduling and scaling</p></li><li><p><strong>Lambda</strong>: Serverless functions. Pay only for compute time consumed. Great for event-driven workloads</p></li><li><p><strong>Route 53</strong>: DNS service with health checks and routing policies</p></li><li><p><strong>ACM</strong>: Free SSL/TLS certificates. Attach to load balancers or CloudFront</p></li><li><p><strong>Secrets Manager</strong>: Stores passwords, API keys, tokens with automatic rotation</p></li></ul></blockquote><h5><strong>AWS Free Tier and surprise bills</strong></h5><p>Three free tier types:</p><blockquote><ul><li><p><strong>12-month free</strong>: 750 hours/month t2.micro/t3.micro EC2, 5GB S3, 750 hours/month RDS single-AZ</p></li><li><p><strong>Always free</strong>: 1M Lambda requests/month, 25GB DynamoDB, 1M SNS notifications</p></li><li><p><strong>Short-term trials</strong>: Service-specific trials (e.g., 750 hours Redshift for 2 months)</p></li></ul></blockquote><p><strong>How to avoid surprise bills:</strong></p><blockquote><ul><li><p><strong>Set up billing alerts</strong> in CloudWatch and <strong>AWS Budgets</strong> with notifications at 50%, 80%, 100%</p></li><li><p><strong>Check billing dashboard</strong> weekly while learning</p></li><li><p><strong>Terminate unused resources</strong>: stopped EC2 instances still incur EBS charges</p></li><li><p><strong>Watch NAT gateways</strong>: ~$32/month even with zero traffic</p></li><li><p><strong>Watch data transfer</strong>: outbound data from AWS costs money</p></li></ul></blockquote><p>Set up a billing alarm with the CLI:</p><pre><code># Create an SNS topic for billing alerts
aws sns create-topic --name billing-alerts --profile dev

# Subscribe your email
aws sns subscribe \
  --topic-arn arn:aws:sns:us-east-1:123456789012:billing-alerts \
  --protocol email \
  --notification-endpoint your-email@example.com \
  --profile dev

# Create CloudWatch alarm for charges over $10
aws cloudwatch put-metric-alarm \
  --alarm-name "billing-alarm-10-usd" \
  --alarm-description "Alarm when charges exceed $10" \
  --metric-name EstimatedCharges \
  --namespace AWS/Billing \
  --statistic Maximum \
  --period 21600 \
  --threshold 10 \
  --comparison-operator GreaterThanThreshold \
  --dimensions Name=Currency,Value=USD \
  --evaluation-periods 1 \
  --alarm-actions arn:aws:sns:us-east-1:123456789012:billing-alerts \
  --region us-east-1 \
  --profile dev</code></pre><p>Billing metrics are only available in <code>us-east-1</code>, regardless of where your resources live.</p><h5><strong>The AWS Well-Architected Framework</strong></h5><p>AWS defines six pillars for well-designed cloud systems:</p><blockquote><ul><li><p><strong>Operational Excellence</strong>: Automate operations, respond to events, define standards</p></li><li><p><strong>Security</strong>: Protect data with IAM, encryption, and detective controls</p></li><li><p><strong>Reliability</strong>: Recover from failures, meet demand, test recovery procedures</p></li><li><p><strong>Performance Efficiency</strong>: Right-size instances, choose correct storage, monitor performance</p></li><li><p><strong>Cost Optimization</strong>: Avoid waste with reserved/spot instances and right-sizing</p></li><li><p><strong>Sustainability</strong>: Minimize environmental impact, optimize utilization</p></li></ul></blockquote><p>We will apply these principles as we build real infrastructure in upcoming articles.</p><h5><strong>Closing notes</strong></h5><p>AWS can feel overwhelming, but most apps only use a handful of services. Once you understand IAM, VPC, and the core compute and storage services, you have the foundation for everything else. Key takeaways: never use root, always apply least privilege, understand public vs private subnets, and set up billing alerts before you forget. In the next article, we will start building real infrastructure with Terraform, defining VPCs, subnets, security groups, and EC2 instances as code.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Your First CI Pipeline with GitHub Actions]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-your-first-ci-pipeline</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-your-first-ci-pipeline</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Sun, 03 May 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0_11!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0_11!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0_11!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!0_11!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!0_11!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!0_11!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0_11!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043145?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!0_11!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!0_11!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!0_11!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!0_11!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc95cce6-4f7b-4378-8231-b5d781a68a2c_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article five of the DevOps from Zero to Hero series. In the previous article we wrote unit and integration tests for a TypeScript project. Tests are great, but they only help if someone actually runs them. That someone should not be you, manually, right before a deploy. It should be a machine that runs them every single time code changes.</p><p>That is what Continuous Integration (CI) is about: automating the boring, repetitive, critical stuff so humans can focus on writing code. In this article we are going to build a complete CI pipeline with GitHub Actions from scratch. By the end, every push and pull request to your repository will automatically lint the code, run tests, build a Docker image, and push it to a container registry.</p><p>Let&#8217;s get into it.</p><h5><strong>What is CI and why it matters</strong></h5><p>Continuous Integration is the practice of automatically building and testing code every time someone pushes a change. The word &#8220;continuous&#8221; is important: this is not something you do once a week or before a release. It happens on every commit, every pull request, every time.</p><p>Why does this matter? Three reasons:</p><blockquote><ul><li><p><strong>Catch bugs early</strong>: A bug found in CI costs minutes to fix. A bug found in production costs hours, customer trust, and sometimes money. The earlier you catch it, the cheaper it is.</p></li><li><p><strong>Enforce standards</strong>: Linting, formatting, and type checking should not depend on developers remembering to run them. CI enforces these standards automatically, every time.</p></li><li><p><strong>Automate repetitive tasks</strong>: Building Docker images, running test suites, generating artifacts. These are things a machine should do, not a person.</p></li></ul></blockquote><p>Without CI, your workflow looks like this: a developer writes code, forgets to run the linter, pushes to main, breaks the build, and the whole team notices an hour later. With CI, the linter runs automatically, the push is blocked, and the developer fixes it in five minutes before anyone else is affected.</p><p>CI is the first real automation layer in a DevOps pipeline. Everything else, continuous delivery, continuous deployment, infrastructure as code, all of it builds on top of CI.</p><h5><strong>GitHub Actions fundamentals</strong></h5><p>GitHub Actions is a CI/CD platform built into GitHub. You define workflows as YAML files in a <code>.github/workflows/</code> directory, and GitHub runs them for you on hosted virtual machines. There is no separate service to set up, no webhooks to configure, and no servers to manage.</p><p>Before we write any YAML, let&#8217;s understand the key concepts:</p><blockquote><ul><li><p><strong>Workflow</strong>: A YAML file that defines an automated process. Each workflow lives in <code>.github/workflows/</code> and is triggered by events.</p></li><li><p><strong>Event (trigger)</strong>: What causes the workflow to run. Common triggers are <code>push</code>, <code>pull_request</code>, and <code>schedule</code>.</p></li><li><p><strong>Job</strong>: A set of steps that run on the same virtual machine (called a &#8220;runner&#8221;). A workflow can have multiple jobs, and by default they run in parallel.</p></li><li><p><strong>Step</strong>: A single task within a job. A step can run a shell command or use a pre-built action.</p></li><li><p><strong>Action</strong>: A reusable unit of code that performs a common task. For example, <code>actions/checkout@v4</code> clones your repository, and <code>actions/setup-node@v4</code> installs Node.js.</p></li><li><p><strong>Runner</strong>: The virtual machine that executes your job. GitHub provides hosted runners with Ubuntu, Windows, and macOS.</p></li></ul></blockquote><p>Here is the hierarchy visualized:</p><pre><code>Workflow (.github/workflows/ci.yml)
  &#9500;&#9472;&#9472; Event: push to main, pull_request
  &#9500;&#9472;&#9472; Job: lint
  &#9474;     &#9500;&#9472;&#9472; Step: Checkout code
  &#9474;     &#9500;&#9472;&#9472; Step: Setup Node.js
  &#9474;     &#9492;&#9472;&#9472; Step: Run ESLint
  &#9500;&#9472;&#9472; Job: test
  &#9474;     &#9500;&#9472;&#9472; Step: Checkout code
  &#9474;     &#9500;&#9472;&#9472; Step: Setup Node.js
  &#9474;     &#9500;&#9472;&#9472; Step: Install dependencies
  &#9474;     &#9492;&#9472;&#9472; Step: Run Vitest
  &#9492;&#9472;&#9472; Job: build
        &#9500;&#9472;&#9472; Step: Checkout code
        &#9500;&#9472;&#9472; Step: Setup Docker Buildx
        &#9492;&#9472;&#9472; Step: Build and push image</code></pre><h5><strong>Triggers: when does CI run?</strong></h5><p>The <code>on</code> key in your workflow file defines when it runs. Here are the triggers you will use most often:</p><pre><code># Run on every push to main
on:
  push:
    branches: [main]

# Run on every pull request targeting main
on:
  pull_request:
    branches: [main]

# Run on both push and pull request
on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

# Run on a schedule (cron syntax, every day at 6 AM UTC)
on:
  schedule:
    - cron: "0 6 * * *"

# Run manually from the GitHub UI
on:
  workflow_dispatch:</code></pre><p>For a typical project, you want CI to run on both <code>push</code> and <code>pull_request</code> to the main branch. The push trigger catches anything that lands on main directly, and the pull request trigger gives you feedback before merging.</p><h5><strong>Building the pipeline step by step</strong></h5><p>Let&#8217;s build a real CI pipeline for a TypeScript project. We will start simple and add features incrementally. Create the file <code>.github/workflows/ci.yml</code> in your repository:</p><pre><code># .github/workflows/ci.yml
name: CI

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

jobs:
  lint:
    name: Lint
    runs-on: ubuntu-latest

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Setup Node.js
        uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"

      - name: Install dependencies
        run: npm ci

      - name: Run ESLint
        run: npx eslint . --max-warnings 0

  test:
    name: Test
    runs-on: ubuntu-latest

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Setup Node.js
        uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"

      - name: Install dependencies
        run: npm ci

      - name: Run tests
        run: npm test

      - name: Run tests with coverage
        run: npm run test:coverage

      - name: Upload coverage report
        uses: actions/upload-artifact@v4
        with:
          name: coverage-report
          path: coverage/</code></pre><p>Let&#8217;s break down what is happening here:</p><blockquote><ul><li><p><strong><code>actions/checkout@v4</code></strong>: Clones your repository into the runner. Without this, the runner has no code to work with.</p></li><li><p><strong><code>actions/setup-node@v4</code></strong>: Installs the specified Node.js version and configures npm caching.</p></li><li><p><strong><code>npm ci</code></strong>: Installs dependencies from <code>package-lock.json</code> exactly as specified. Unlike <code>npm install</code>, it does not modify the lockfile and is faster and more reliable in CI.</p></li><li><p><strong><code>npx eslint . --max-warnings 0</code></strong>: Runs ESLint and fails if there are any warnings. This is stricter than the default, which only fails on errors. Treating warnings as errors in CI prevents them from piling up.</p></li><li><p><strong>Lint and test jobs run in parallel</strong>: Since they do not depend on each other, GitHub runs them at the same time, making your pipeline faster.</p></li></ul></blockquote><h5><strong>Adding the Docker build</strong></h5><p>Now let&#8217;s add a job that builds a Docker image and pushes it to GitHub Container Registry (GHCR). This job should only run after linting and tests pass, so we use the <code>needs</code> keyword to create a dependency:</p><pre><code>  build:
    name: Build and Push Docker Image
    runs-on: ubuntu-latest
    needs: [lint, test]
    if: github.event_name == 'push' &amp;&amp; github.ref == 'refs/heads/main'

    permissions:
      contents: read
      packages: write

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Set up Docker Buildx
        uses: docker/setup-buildx-action@v3

      - name: Log in to GitHub Container Registry
        uses: docker/login-action@v3
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - name: Extract metadata
        id: meta
        uses: docker/metadata-action@v5
        with:
          images: ghcr.io/${{ github.repository }}
          tags: |
            type=sha,prefix=
            type=raw,value=latest

      - name: Build and push
        uses: docker/build-push-action@v6
        with:
          context: .
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          cache-from: type=gha
          cache-to: type=gha,mode=max</code></pre><p>There is a lot going on here, so let&#8217;s unpack it:</p><blockquote><ul><li><p><strong><code>needs: [lint, test]</code></strong>: This job waits for both lint and test to pass before running. If either fails, the build is skipped entirely.</p></li><li><p><strong><code>if: github.event_name == 'push' &amp;&amp; github.ref == 'refs/heads/main'</code></strong>: Only build images on pushes to main, not on pull requests. You do not want to push a Docker image for every PR.</p></li><li><p><strong><code>permissions</code></strong>: GitHub Actions uses a <code>GITHUB_TOKEN</code> that is automatically created for each workflow run. We need <code>packages: write</code> to push to GHCR.</p></li><li><p><strong><code>docker/setup-buildx-action@v3</code></strong>: Sets up Docker Buildx, which is an extended build tool that supports advanced features like caching and multi-platform builds.</p></li><li><p><strong><code>docker/login-action@v3</code></strong>: Logs into GHCR using the built-in <code>GITHUB_TOKEN</code>. No need to create a personal access token.</p></li><li><p><strong><code>docker/metadata-action@v5</code></strong>: Generates tags and labels automatically. We tag with both the Git SHA (for traceability) and <code>latest</code> (for convenience).</p></li><li><p><strong><code>docker/build-push-action@v6</code></strong>: Builds the Dockerfile and pushes the image. The <code>cache-from</code> and <code>cache-to</code> lines enable GitHub Actions cache for Docker layers, which we will explain next.</p></li></ul></blockquote><h5><strong>Caching: making CI fast</strong></h5><p>CI pipelines that take 10 minutes quickly become a bottleneck. Developers stop waiting for them, start merging without checking results, and the whole point of CI breaks down. Caching is how you keep things fast.</p><p>There are two things worth caching in a Node.js project: npm packages and Docker layers.</p><p><strong>npm cache</strong> is the easier one. The <code>actions/setup-node@v4</code> action handles it for you when you add <code>cache: "npm"</code>:</p><pre><code>      - name: Setup Node.js
        uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"</code></pre><p>This caches the npm download cache (not <code>node_modules</code>), so <code>npm ci</code> still runs but does not need to download packages from the registry. The first run populates the cache, and subsequent runs reuse it. On a project with many dependencies, this can save 30 to 60 seconds per run.</p><p><strong>Docker layer cache</strong> is more impactful. Building a Docker image from scratch every time is wasteful because most layers (like the base image and installed system packages) rarely change. Docker Buildx with the GitHub Actions cache backend stores layers between runs:</p><pre><code>      - name: Build and push
        uses: docker/build-push-action@v6
        with:
          context: .
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          cache-from: type=gha
          cache-to: type=gha,mode=max</code></pre><blockquote><ul><li><p><strong><code>cache-from: type=gha</code></strong>: Pull cached layers from the GitHub Actions cache.</p></li><li><p><strong><code>cache-to: type=gha,mode=max</code></strong>: Push all layers to the cache after building. The <code>mode=max</code> option caches intermediate layers too, not just the final image layers.</p></li></ul></blockquote><p>A well-structured Dockerfile benefits enormously from this. If your dependency installation layer has not changed, Docker reuses the cached layer instead of running <code>npm ci</code> again inside the container. This can cut build times from minutes to seconds.</p><h5><strong>Matrix builds: testing across versions</strong></h5><p>Sometimes you need to test your code against multiple Node.js versions, or multiple operating systems, or both. Matrix builds let you define a set of variables and run the job once for each combination.</p><pre><code>  test:
    name: Test (Node ${{ matrix.node-version }})
    runs-on: ubuntu-latest

    strategy:
      matrix:
        node-version: ["20", "22"]
      fail-fast: false

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Setup Node.js ${{ matrix.node-version }}
        uses: actions/setup-node@v4
        with:
          node-version: ${{ matrix.node-version }}
          cache: "npm"

      - name: Install dependencies
        run: npm ci

      - name: Run tests
        run: npm test</code></pre><p>This runs the test job twice: once with Node 20 and once with Node 22. Both runs happen in parallel on separate runners, so it does not slow down your pipeline.</p><p>Key settings:</p><blockquote><ul><li><p><strong><code>strategy.matrix</code></strong>: Defines the variables and their values. You can add more dimensions, like <code>os: [ubuntu-latest, windows-latest]</code>, and GitHub will run every combination.</p></li><li><p><strong><code>fail-fast: false</code></strong>: By default, if one matrix job fails, GitHub cancels the others. Setting this to <code>false</code> lets all jobs complete, so you can see all failures at once.</p></li></ul></blockquote><p>Matrix builds are especially useful for libraries that need to support multiple runtimes. For application code where you control the runtime, testing a single version is usually enough.</p><h5><strong>Secrets and environment variables</strong></h5><p>Your CI pipeline will often need credentials: API keys for external services, tokens for registries, or database passwords for integration tests. GitHub provides two mechanisms for this.</p><p><strong>Environment variables</strong> are for non-sensitive values:</p><pre><code>    env:
      NODE_ENV: test
      API_URL: https://api.staging.example.com

    steps:
      - name: Run tests
        run: npm test
        env:
          DATABASE_URL: postgres://localhost:5432/testdb</code></pre><p>You can set environment variables at the workflow level, job level, or step level. Step-level variables override job-level variables, which override workflow-level variables.</p><p><strong>Secrets</strong> are for sensitive values like API keys and tokens:</p><pre><code>      - name: Deploy to staging
        run: ./deploy.sh
        env:
          DEPLOY_TOKEN: ${{ secrets.DEPLOY_TOKEN }}
          AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
          AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}</code></pre><p>To add secrets, go to your repository&#8217;s Settings, then Secrets and variables, then Actions. Secrets are encrypted at rest and masked in logs. GitHub will replace the secret value with <code>***</code> if it accidentally appears in the output.</p><p>Important rules about secrets:</p><blockquote><ul><li><p><strong>Never hardcode secrets in your workflow files</strong>. They are committed to the repository and visible to anyone with read access.</p></li><li><p><strong><code>GITHUB_TOKEN</code> is automatic</strong>. You do not need to create it. GitHub generates one for every workflow run with permissions scoped to the repository.</p></li><li><p><strong>Secrets are not available in pull requests from forks</strong>. This is a security feature. If your tests need secrets, they will fail on fork PRs, which is expected.</p></li><li><p><strong>Use environments for deployment secrets</strong>. GitHub environments let you require approvals and restrict which branches can use certain secrets.</p></li></ul></blockquote><h5><strong>Reusable workflows: keeping things DRY</strong></h5><p>As your organization grows, you will have multiple repositories that need similar CI pipelines. Copy pasting YAML files between repositories is a maintenance nightmare. Reusable workflows let you define a workflow once and call it from other workflows.</p><p>First, create the reusable workflow in a shared repository. The key difference is the <code>workflow_call</code> trigger:</p><pre><code># .github/workflows/node-ci.yml (in your shared repo)
name: Node.js CI

on:
  workflow_call:
    inputs:
      node-version:
        description: "Node.js version to use"
        required: false
        type: string
        default: "22"
      run-lint:
        description: "Whether to run linting"
        required: false
        type: boolean
        default: true

jobs:
  lint:
    name: Lint
    runs-on: ubuntu-latest
    if: ${{ inputs.run-lint }}

    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: ${{ inputs.node-version }}
          cache: "npm"

      - run: npm ci

      - run: npx eslint . --max-warnings 0

  test:
    name: Test
    runs-on: ubuntu-latest

    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: ${{ inputs.node-version }}
          cache: "npm"

      - run: npm ci

      - run: npm test</code></pre><p>Then call it from any repository:</p><pre><code># .github/workflows/ci.yml (in your project repo)
name: CI

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

jobs:
  ci:
    uses: your-org/shared-workflows/.github/workflows/node-ci.yml@main
    with:
      node-version: "22"
      run-lint: true</code></pre><p>The benefits are significant:</p><blockquote><ul><li><p><strong>Single source of truth</strong>: Update the shared workflow and every repository that uses it gets the update.</p></li><li><p><strong>Consistency</strong>: Every project follows the same CI process, same actions versions, same caching strategy.</p></li><li><p><strong>Less maintenance</strong>: Fix a bug or upgrade an action in one place, not in fifty repositories.</p></li><li><p><strong>Inputs make it flexible</strong>: Each project can customize behavior (Node version, whether to lint, etc.) without forking the workflow.</p></li></ul></blockquote><h5><strong>Status badges: show your pipeline health</strong></h5><p>Once your CI pipeline is working, you want everyone to see its status at a glance. GitHub provides status badges that you can add to your README:</p><pre><code>![CI](https://github.com/your-org/your-repo/actions/workflows/ci.yml/badge.svg)</code></pre><p>This renders as a small badge that shows &#8220;passing&#8221; (green) or &#8220;failing&#8221; (red) based on the latest run of the workflow. Add it to the top of your README so contributors immediately know the project&#8217;s health.</p><p>You can also make badges branch-specific:</p><pre><code>![CI](https://github.com/your-org/your-repo/actions/workflows/ci.yml/badge.svg?branch=main)</code></pre><p>This only reflects the status of the workflow on the main branch, ignoring feature branches.</p><h5><strong>Branch protection: require CI to pass before merge</strong></h5><p>A CI pipeline is only useful if people cannot bypass it. Branch protection rules ensure that code cannot be merged into main unless CI passes. Here is how to set it up:</p><blockquote><ol><li><p>Go to your repository&#8217;s Settings, then Branches.</p></li><li><p>Click &#8220;Add branch protection rule&#8221; (or &#8220;Add classic branch protection rule&#8221;).</p></li><li><p>Set the branch name pattern to <code>main</code>.</p></li><li><p>Check &#8220;Require status checks to pass before merging.&#8221;</p></li><li><p>Search for and select your CI job names (e.g., &#8220;Lint&#8221;, &#8220;Test&#8221;).</p></li><li><p>Optionally check &#8220;Require branches to be up to date before merging&#8221; to prevent merging stale branches.</p></li></ol></blockquote><p>With this in place, the merge button on a pull request is disabled until all required checks pass. No one can bypass CI, not even repository admins (unless they explicitly override it, which leaves an audit trail).</p><p>Additional protections worth enabling:</p><blockquote><ul><li><p><strong>Require pull request reviews</strong>: At least one team member must approve before merging.</p></li><li><p><strong>Require linear history</strong>: Force squash or rebase merges for a clean git history.</p></li><li><p><strong>Do not allow bypassing the above settings</strong>: Even admins must follow the rules.</p></li></ul></blockquote><h5><strong>The complete workflow file</strong></h5><p>Here is the full CI pipeline combining everything we covered. This is a production-ready starting point for any TypeScript project:</p><pre><code># .github/workflows/ci.yml
name: CI

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

env:
  NODE_ENV: test
  REGISTRY: ghcr.io
  IMAGE_NAME: ${{ github.repository }}

jobs:
  lint:
    name: Lint
    runs-on: ubuntu-latest

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Setup Node.js
        uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"

      - name: Install dependencies
        run: npm ci

      - name: Check formatting
        run: npx prettier --check .

      - name: Run ESLint
        run: npx eslint . --max-warnings 0

      - name: Type check
        run: npx tsc --noEmit

  test:
    name: Test (Node ${{ matrix.node-version }})
    runs-on: ubuntu-latest

    strategy:
      matrix:
        node-version: ["20", "22"]
      fail-fast: false

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Setup Node.js ${{ matrix.node-version }}
        uses: actions/setup-node@v4
        with:
          node-version: ${{ matrix.node-version }}
          cache: "npm"

      - name: Install dependencies
        run: npm ci

      - name: Run tests
        run: npm test

      - name: Run tests with coverage
        run: npm run test:coverage

      - name: Upload coverage report
        if: matrix.node-version == '22'
        uses: actions/upload-artifact@v4
        with:
          name: coverage-report
          path: coverage/
          retention-days: 14

  build:
    name: Build and Push Docker Image
    runs-on: ubuntu-latest
    needs: [lint, test]
    if: github.event_name == 'push' &amp;&amp; github.ref == 'refs/heads/main'

    permissions:
      contents: read
      packages: write

    steps:
      - name: Checkout code
        uses: actions/checkout@v4

      - name: Set up Docker Buildx
        uses: docker/setup-buildx-action@v3

      - name: Log in to GHCR
        uses: docker/login-action@v3
        with:
          registry: ${{ env.REGISTRY }}
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}

      - name: Extract metadata
        id: meta
        uses: docker/metadata-action@v5
        with:
          images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}
          tags: |
            type=sha,prefix=
            type=raw,value=latest

      - name: Build and push
        uses: docker/build-push-action@v6
        with:
          context: .
          push: true
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
          cache-from: type=gha
          cache-to: type=gha,mode=max</code></pre><p>Notice a few things about this complete workflow:</p><blockquote><ul><li><p><strong>Three stages</strong>: Lint, test, and build. They form a pipeline where each stage gates the next.</p></li><li><p><strong>Type checking in lint</strong>: We added <code>tsc --noEmit</code> to catch TypeScript errors. This is a cheap check that catches a whole class of bugs.</p></li><li><p><strong>Prettier check</strong>: <code>prettier --check</code> verifies formatting without modifying files. If a developer forgot to format, CI catches it.</p></li><li><p><strong>Coverage only uploaded once</strong>: When running a matrix build, you only need one coverage report, not one per Node version. The <code>if: matrix.node-version == '22'</code> conditional handles this.</p></li><li><p><strong>Retention days</strong>: Artifacts do not need to live forever. Setting <code>retention-days: 14</code> keeps things tidy.</p></li><li><p><strong>Environment variables at the top</strong>: <code>REGISTRY</code> and <code>IMAGE_NAME</code> are defined once and reused, making the workflow easier to adapt to other registries.</p></li></ul></blockquote><h5><strong>Debugging failed workflows</strong></h5><p>When your CI pipeline fails (and it will), here is how to debug it:</p><blockquote><ul><li><p><strong>Read the logs</strong>: Click on the failed job in the GitHub Actions UI. Each step shows its output. The error is usually in the last few lines of the failed step.</p></li><li><p><strong>Run locally first</strong>: Before pushing, run the same commands locally. <code>npm ci &amp;&amp; npx eslint . &amp;&amp; npm test</code> should produce the same result as CI.</p></li><li><p><strong>Check the runner environment</strong>: CI runs on a clean Ubuntu machine. If something works locally but fails in CI, the difference is usually in environment variables, installed tools, or file paths.</p></li><li><p><strong>Use <code>act</code> for local testing</strong>: The <code>act</code> tool (<a href="https://github.com/nektos/act">https://github.com/nektos/act</a>) lets you run GitHub Actions workflows on your local machine using Docker. It is not perfect, but it catches most issues.</p></li><li><p><strong>Enable debug logging</strong>: Re-run the workflow with debug logging enabled by going to the failed run, clicking &#8220;Re-run all jobs&#8221;, and checking &#8220;Enable debug logging.&#8221; This adds verbose output from every action.</p></li></ul></blockquote><h5><strong>Common pitfalls and how to avoid them</strong></h5><p>A few things that trip people up when setting up CI for the first time:</p><blockquote><ul><li><p><strong>Not using <code>npm ci</code></strong>: Using <code>npm install</code> in CI can produce different dependency trees than your local machine. Always use <code>npm ci</code>, which installs exactly what is in <code>package-lock.json</code>.</p></li><li><p><strong>Missing <code>package-lock.json</code> in the repository</strong>: If you gitignored it, <code>npm ci</code> will fail. The lockfile should always be committed.</p></li><li><p><strong>Tests that depend on order</strong>: If your tests pass locally but fail in CI, they might depend on execution order. Vitest runs tests in parallel by default, which can expose this.</p></li><li><p><strong>Hardcoded paths</strong>: Tests that reference <code>/Users/yourname/project/</code> will fail on a Linux runner. Use relative paths or environment variables.</p></li><li><p><strong>Forgetting the Docker context</strong>: If your Dockerfile copies files with <code>COPY . .</code>, make sure your <code>.dockerignore</code> excludes <code>node_modules</code>, <code>.git</code>, and other large directories.</p></li><li><p><strong>Overly broad triggers</strong>: Running CI on every push to every branch wastes runner minutes. Limit triggers to <code>main</code> and pull requests targeting <code>main</code>.</p></li></ul></blockquote><h5><strong>What comes next</strong></h5><p>We now have a CI pipeline that lints, tests, and builds our code automatically. But CI is only half the story. Getting code into a container is useful, but that container needs to go somewhere.</p><p>In the next article, we will tackle Continuous Deployment (CD): taking the Docker image we just built and deploying it to a real environment. We will cover deployment strategies, rollbacks, and how to make deployments boring (which is exactly what you want them to be).</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Automated Testing]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-automated-testing</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-automated-testing</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Thu, 30 Apr 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nZhQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nZhQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nZhQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!nZhQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!nZhQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!nZhQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nZhQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043147?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nZhQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!nZhQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!nZhQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!nZhQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F443ac2f6-4122-46ee-99c5-0efe1ccac8de_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article four of the DevOps from Zero to Hero series. In the previous articles we covered the fundamentals of Linux, networking, and version control with Git. Now it is time to talk about something that separates hobby projects from production-ready software: automated testing.</p><p>If you have ever pushed a change to production and immediately regretted it, you already understand why testing matters. Automated tests give you confidence that your code works as expected before it reaches users. In a DevOps context, tests are the gate between &#8220;code written&#8221; and &#8220;code deployed.&#8221; Without them, your CI/CD pipeline is just a fast way to ship bugs.</p><p>In this article we will cover the testing pyramid, write real unit and integration tests in TypeScript using Vitest and Supertest, talk about what coverage actually means (and why chasing 100% is a trap), and lay the groundwork for running tests in CI, which we will cover in depth in article five.</p><p>Let&#8217;s get into it.</p><h5><strong>Why testing matters for DevOps</strong></h5><p>Testing is not just a developer concern. In a DevOps workflow, tests are the foundation of everything else you build. Here is why:</p><blockquote><ul><li><p><strong>Confidence to deploy</strong>: If your tests pass, you can deploy without fear. If they do not, you know something is broken before users do.</p></li><li><p><strong>Fast feedback</strong>: A good test suite tells you within minutes whether a change is safe. Compare that to waiting for manual QA or finding out from a user report.</p></li><li><p><strong>Catch regressions</strong>: Code that worked yesterday can break today because of a seemingly unrelated change. Tests catch these regressions automatically.</p></li><li><p><strong>Enable automation</strong>: CI/CD pipelines depend on tests. Without automated tests, your pipeline is just automated deployment of untested code.</p></li><li><p><strong>Documentation</strong>: Well-written tests describe what your code should do. They serve as living documentation that stays in sync with the actual behavior.</p></li></ul></blockquote><p>Think of it this way: every test you write is a tiny contract that says &#8220;this behavior must be preserved.&#8221; When someone changes the code six months from now, those contracts catch anything that breaks. That is incredibly valuable in a team environment where multiple people touch the same codebase.</p><h5><strong>The testing pyramid</strong></h5><p>The testing pyramid is a model that helps you decide how many tests of each type to write. It looks like this:</p><pre><code>        /  E2E  \          Few, slow, expensive
       /----------\
      / Integration \      Some, moderate speed
     /----------------\
    /    Unit Tests     \  Many, fast, cheap
   /____________________\</code></pre><p>The shape matters. Here is why:</p><blockquote><ul><li><p><strong>Unit tests</strong> (base of the pyramid): These test individual functions or modules in isolation. They are fast, cheap to write, and cheap to run. You should have the most of these.</p></li><li><p><strong>Integration tests</strong> (middle): These test how multiple pieces work together, like an API endpoint hitting a database. They are slower and more complex, but they catch issues that unit tests miss.</p></li><li><p><strong>End-to-end tests</strong> (top): These test the entire application from the user&#8217;s perspective, often through a browser. They are the slowest, most fragile, and most expensive to maintain. You should have the fewest of these.</p></li></ul></blockquote><p>The pyramid shape exists because of a tradeoff between speed and confidence. Unit tests run in milliseconds but only test small pieces. E2E tests take seconds or minutes but test the full flow. If you invert the pyramid (lots of E2E, few unit tests), your test suite becomes slow, flaky, and painful to maintain.</p><p>A healthy ratio might look something like 70% unit, 20% integration, 10% E2E. These numbers are not rules, they are guidelines. The key insight is: push testing down to the lowest level that gives you confidence. If you can catch a bug with a unit test, do not write an E2E test for it.</p><h5><strong>Setting up the project</strong></h5><p>Let&#8217;s build a small TypeScript project with tests. We will use Vitest as our test runner because it is fast, modern, and works great with TypeScript out of the box.</p><p>First, initialize the project:</p><pre><code>mkdir testing-demo &amp;&amp; cd testing-demo
npm init -y
npm install -D typescript vitest @types/node
npm install express
npm install -D @types/express supertest @types/supertest</code></pre><p>Create a <code>tsconfig.json</code>:</p><pre><code>{
  "compilerOptions": {
    "target": "ES2022",
    "module": "ESNext",
    "moduleResolution": "bundler",
    "strict": true,
    "esModuleInterop": true,
    "outDir": "./dist",
    "rootDir": "./src",
    "declaration": true,
    "sourceMap": true
  },
  "include": ["src/**/*"],
  "exclude": ["node_modules", "dist"]
}</code></pre><p>Add the test script to <code>package.json</code>:</p><pre><code>{
  "scripts": {
    "test": "vitest run",
    "test:watch": "vitest",
    "test:coverage": "vitest run --coverage"
  }
}</code></pre><h5><strong>Unit testing with Vitest</strong></h5><p>Let&#8217;s start with the base of the pyramid. Unit tests verify that individual functions do what they are supposed to do. They should be fast, isolated, and deterministic.</p><p>Here is a simple utility module at <code>src/utils.ts</code>:</p><pre><code>// src/utils.ts

export function slugify(text: string): string {
  return text
    .toLowerCase()
    .trim()
    .replace(/[^\w\s-]/g, "")
    .replace(/[\s_]+/g, "-")
    .replace(/-+/g, "-")
    .replace(/^-|-$/g, "");
}

export function truncate(text: string, maxLength: number): string {
  if (text.length &lt;= maxLength) {
    return text;
  }
  const truncated = text.slice(0, maxLength);
  const lastSpace = truncated.lastIndexOf(" ");
  if (lastSpace &gt; 0) {
    return truncated.slice(0, lastSpace) + "...";
  }
  return truncated + "...";
}

export function parseQueryParams(query: string): Record&lt;string, string&gt; {
  if (!query || query.trim() === "") {
    return {};
  }
  const cleaned = query.startsWith("?") ? query.slice(1) : query;
  return cleaned.split("&amp;").reduce(
    (params, pair) =&gt; {
      const [key, value] = pair.split("=");
      if (key) {
        params[decodeURIComponent(key)] = decodeURIComponent(value ?? "");
      }
      return params;
    },
    {} as Record&lt;string, string&gt;,
  );
}</code></pre><p>Now let&#8217;s write the tests at <code>src/utils.test.ts</code>:</p><pre><code>// src/utils.test.ts
import { describe, it, expect } from "vitest";
import { slugify, truncate, parseQueryParams } from "./utils";

describe("slugify", () =&gt; {
  it("converts a simple string to a slug", () =&gt; {
    expect(slugify("Hello World")).toBe("hello-world");
  });

  it("handles special characters", () =&gt; {
    expect(slugify("Hello, World! How's it going?")).toBe(
      "hello-world-hows-it-going",
    );
  });

  it("collapses multiple spaces and dashes", () =&gt; {
    expect(slugify("too   many   spaces")).toBe("too-many-spaces");
    expect(slugify("too---many---dashes")).toBe("too-many-dashes");
  });

  it("trims leading and trailing dashes", () =&gt; {
    expect(slugify("  -hello-  ")).toBe("hello");
  });

  it("handles empty string", () =&gt; {
    expect(slugify("")).toBe("");
  });
});

describe("truncate", () =&gt; {
  it("returns the full string if it is shorter than maxLength", () =&gt; {
    expect(truncate("short", 10)).toBe("short");
  });

  it("returns the full string if it equals maxLength", () =&gt; {
    expect(truncate("exact", 5)).toBe("exact");
  });

  it("truncates at the last space before maxLength", () =&gt; {
    expect(truncate("this is a longer sentence", 15)).toBe("this is a...");
  });

  it("truncates without space if no space is found", () =&gt; {
    expect(truncate("superlongwordwithoutspaces", 10)).toBe(
      "superlongw...",
    );
  });
});

describe("parseQueryParams", () =&gt; {
  it("parses a simple query string", () =&gt; {
    expect(parseQueryParams("name=alice&amp;age=30")).toEqual({
      name: "alice",
      age: "30",
    });
  });

  it("handles a leading question mark", () =&gt; {
    expect(parseQueryParams("?foo=bar")).toEqual({ foo: "bar" });
  });

  it("handles URL-encoded values", () =&gt; {
    expect(parseQueryParams("msg=hello%20world")).toEqual({
      msg: "hello world",
    });
  });

  it("returns an empty object for empty input", () =&gt; {
    expect(parseQueryParams("")).toEqual({});
  });

  it("handles keys without values", () =&gt; {
    expect(parseQueryParams("flag=")).toEqual({ flag: "" });
  });
});</code></pre><p>Let&#8217;s break down the test structure:</p><blockquote><ul><li><p><strong><code>describe</code></strong> groups related tests. Think of it as a section header for a set of behaviors.</p></li><li><p><strong><code>it</code></strong> defines an individual test case. The string should read like a sentence: &#8220;it converts a simple string to a slug.&#8221;</p></li><li><p><strong><code>expect</code></strong> is the assertion. It takes a value and chains a matcher like <code>toBe</code>, <code>toEqual</code>, <code>toContain</code>, or <code>toThrow</code>.</p></li></ul></blockquote><p>Run the tests:</p><pre><code>npx vitest run

# Output:
# &#10003; src/utils.test.ts (10 tests) 5ms
#   &#10003; slugify (5 tests)
#   &#10003; truncate (4 tests)
#   &#10003; parseQueryParams (5 tests)
# Test Files  1 passed (1)
# Tests       14 passed (14)</code></pre><h5><strong>Mocking dependencies</strong></h5><p>Real-world code has dependencies: databases, APIs, file systems. In unit tests, you want to isolate the function under test by replacing those dependencies with controlled substitutes. This is called mocking.</p><p>Here is a module that depends on an external API at <code>src/weather.ts</code>:</p><pre><code>// src/weather.ts

export interface WeatherData {
  city: string;
  temperature: number;
  description: string;
}

export async function fetchWeather(city: string): Promise&lt;WeatherData&gt; {
  const response = await fetch(
    `https://api.weather.example.com/v1/current?city=${encodeURIComponent(city)}`,
  );
  if (!response.ok) {
    throw new Error(`Weather API returned ${response.status}`);
  }
  const data = await response.json();
  return {
    city: data.location.name,
    temperature: data.current.temp_c,
    description: data.current.condition.text,
  };
}

export function formatWeatherReport(weather: WeatherData): string {
  return `${weather.city}: ${weather.temperature}C, ${weather.description}`;
}</code></pre><p>And the tests at <code>src/weather.test.ts</code>:</p><pre><code>// src/weather.test.ts
import { describe, it, expect, vi, beforeEach } from "vitest";
import { fetchWeather, formatWeatherReport } from "./weather";

// Mock the global fetch function
const mockFetch = vi.fn();
vi.stubGlobal("fetch", mockFetch);

beforeEach(() =&gt; {
  mockFetch.mockReset();
});

describe("fetchWeather", () =&gt; {
  it("returns parsed weather data on success", async () =&gt; {
    mockFetch.mockResolvedValueOnce({
      ok: true,
      json: async () =&gt; ({
        location: { name: "London" },
        current: { temp_c: 15, condition: { text: "Partly cloudy" } },
      }),
    });

    const result = await fetchWeather("London");

    expect(result).toEqual({
      city: "London",
      temperature: 15,
      description: "Partly cloudy",
    });
    expect(mockFetch).toHaveBeenCalledWith(
      "https://api.weather.example.com/v1/current?city=London",
    );
  });

  it("throws on non-ok response", async () =&gt; {
    mockFetch.mockResolvedValueOnce({
      ok: false,
      status: 404,
    });

    await expect(fetchWeather("Nowhere")).rejects.toThrow(
      "Weather API returned 404",
    );
  });
});

describe("formatWeatherReport", () =&gt; {
  it("formats the weather data as a readable string", () =&gt; {
    const weather = {
      city: "Berlin",
      temperature: 22,
      description: "Sunny",
    };
    expect(formatWeatherReport(weather)).toBe("Berlin: 22C, Sunny");
  });
});</code></pre><p>Key mocking concepts:</p><blockquote><ul><li><p><strong><code>vi.fn()</code></strong> creates a mock function that records how it was called.</p></li><li><p><strong><code>vi.stubGlobal()</code></strong> replaces a global like <code>fetch</code> with your mock.</p></li><li><p><strong><code>mockResolvedValueOnce()</code></strong> tells the mock what to return the next time it is called.</p></li><li><p><strong><code>mockReset()</code></strong> clears the mock state between tests so they do not leak into each other.</p></li></ul></blockquote><p>The important thing to understand about mocking is this: you are not testing <code>fetch</code>. You are testing that your code correctly handles the response from <code>fetch</code>. The mock lets you simulate different scenarios (success, error, timeout) without making real network calls.</p><h5><strong>Integration testing with Supertest</strong></h5><p>Integration tests verify that multiple pieces of your application work together. For web applications, the most common integration test is hitting an API endpoint and verifying the response.</p><p>Here is a simple Express app at <code>src/app.ts</code>:</p><pre><code>// src/app.ts
import express from "express";
import { slugify, truncate } from "./utils";

export const app = express();

app.use(express.json());

interface Article {
  id: number;
  title: string;
  slug: string;
  content: string;
  summary?: string;
}

const articles: Article[] = [];
let nextId = 1;

app.get("/api/articles", (_req, res) =&gt; {
  res.json(articles);
});

app.get("/api/articles/:slug", (req, res) =&gt; {
  const article = articles.find((a) =&gt; a.slug === req.params.slug);
  if (!article) {
    res.status(404).json({ error: "Article not found" });
    return;
  }
  res.json(article);
});

app.post("/api/articles", (req, res) =&gt; {
  const { title, content } = req.body;
  if (!title || !content) {
    res.status(400).json({ error: "Title and content are required" });
    return;
  }
  const article: Article = {
    id: nextId++,
    title,
    slug: slugify(title),
    content,
    summary: truncate(content, 100),
  };
  articles.push(article);
  res.status(201).json(article);
});

app.delete("/api/articles/:slug", (req, res) =&gt; {
  const index = articles.findIndex((a) =&gt; a.slug === req.params.slug);
  if (index === -1) {
    res.status(404).json({ error: "Article not found" });
    return;
  }
  articles.splice(index, 1);
  res.status(204).send();
});</code></pre><p>Now the integration tests at <code>src/app.test.ts</code>:</p><pre><code>// src/app.test.ts
import { describe, it, expect, beforeEach } from "vitest";
import request from "supertest";
import { app } from "./app";

describe("Articles API", () =&gt; {
  // Note: In a real app, you would reset the database between tests.
  // Here we rely on the in-memory array.

  describe("POST /api/articles", () =&gt; {
    it("creates a new article", async () =&gt; {
      const response = await request(app)
        .post("/api/articles")
        .send({
          title: "My First Post",
          content:
            "This is the content of my first blog post. It has enough words to test truncation properly.",
        })
        .expect(201);

      expect(response.body).toMatchObject({
        title: "My First Post",
        slug: "my-first-post",
        content:
          "This is the content of my first blog post. It has enough words to test truncation properly.",
      });
      expect(response.body.id).toBeDefined();
      expect(response.body.summary).toBeDefined();
    });

    it("returns 400 when title is missing", async () =&gt; {
      const response = await request(app)
        .post("/api/articles")
        .send({ content: "some content" })
        .expect(400);

      expect(response.body.error).toBe("Title and content are required");
    });

    it("returns 400 when content is missing", async () =&gt; {
      const response = await request(app)
        .post("/api/articles")
        .send({ title: "A Title" })
        .expect(400);

      expect(response.body.error).toBe("Title and content are required");
    });
  });

  describe("GET /api/articles", () =&gt; {
    it("returns the list of articles", async () =&gt; {
      const response = await request(app).get("/api/articles").expect(200);

      expect(Array.isArray(response.body)).toBe(true);
      expect(response.body.length).toBeGreaterThan(0);
    });
  });

  describe("GET /api/articles/:slug", () =&gt; {
    it("returns an article by slug", async () =&gt; {
      const response = await request(app)
        .get("/api/articles/my-first-post")
        .expect(200);

      expect(response.body.slug).toBe("my-first-post");
      expect(response.body.title).toBe("My First Post");
    });

    it("returns 404 for a non-existent slug", async () =&gt; {
      const response = await request(app)
        .get("/api/articles/does-not-exist")
        .expect(404);

      expect(response.body.error).toBe("Article not found");
    });
  });

  describe("DELETE /api/articles/:slug", () =&gt; {
    it("deletes an article by slug", async () =&gt; {
      // First, create an article to delete
      await request(app)
        .post("/api/articles")
        .send({ title: "To Be Deleted", content: "This will be removed" });

      await request(app)
        .delete("/api/articles/to-be-deleted")
        .expect(204);

      // Verify it is gone
      await request(app)
        .get("/api/articles/to-be-deleted")
        .expect(404);
    });

    it("returns 404 when deleting a non-existent article", async () =&gt; {
      await request(app)
        .delete("/api/articles/ghost-article")
        .expect(404);
    });
  });
});</code></pre><p>Notice how integration tests differ from unit tests:</p><blockquote><ul><li><p><strong>They test the full request/response cycle</strong>, not just a single function.</p></li><li><p><strong>They exercise multiple layers</strong> (routing, validation, business logic) together.</p></li><li><p><strong>They are slower</strong> because they spin up the HTTP layer, but they catch bugs that unit tests cannot, like incorrect route definitions or missing middleware.</p></li></ul></blockquote><p>Supertest is excellent because it does not require you to start the server on a port. It hooks directly into Express, so tests are fast and do not conflict with each other.</p><h5><strong>Testing against real databases with Testcontainers</strong></h5><p>For applications that use a database, you need to decide: do you mock the database or use a real one? Mocking is faster but can hide bugs related to SQL syntax, constraints, or query behavior. Testcontainers gives you the best of both worlds by spinning up a real database in Docker for your tests.</p><p>Here is what using Testcontainers looks like conceptually:</p><pre><code>// src/db.integration.test.ts (conceptual example)
import { describe, it, expect, beforeAll, afterAll } from "vitest";
import { PostgreSqlContainer } from "@testcontainers/postgresql";
import { Client } from "pg";

describe("Database integration", () =&gt; {
  let container: any;
  let client: Client;

  beforeAll(async () =&gt; {
    // Start a real PostgreSQL container
    container = await new PostgreSqlContainer("postgres:16")
      .withDatabase("testdb")
      .start();

    client = new Client({
      connectionString: container.getConnectionUri(),
    });
    await client.connect();

    // Run migrations
    await client.query(`
      CREATE TABLE articles (
        id SERIAL PRIMARY KEY,
        title TEXT NOT NULL,
        slug TEXT UNIQUE NOT NULL,
        content TEXT NOT NULL,
        created_at TIMESTAMPTZ DEFAULT NOW()
      )
    `);
  }, 60000); // Containers can take a moment to start

  afterAll(async () =&gt; {
    await client.end();
    await container.stop();
  });

  it("inserts and retrieves an article", async () =&gt; {
    await client.query(
      "INSERT INTO articles (title, slug, content) VALUES ($1, $2, $3)",
      ["Test Article", "test-article", "Some content"],
    );

    const result = await client.query(
      "SELECT * FROM articles WHERE slug = $1",
      ["test-article"],
    );

    expect(result.rows).toHaveLength(1);
    expect(result.rows[0].title).toBe("Test Article");
  });
});</code></pre><p>Testcontainers is especially useful because:</p><blockquote><ul><li><p><strong>Tests run against the same database engine</strong> you use in production, catching driver-specific bugs.</p></li><li><p><strong>Each test suite gets a fresh container</strong>, so tests do not interfere with each other.</p></li><li><p><strong>It works in CI</strong> as long as Docker is available (which it usually is in GitHub Actions).</p></li><li><p><strong>No shared test database</strong> that accumulates stale data or causes flaky tests from parallel runs.</p></li></ul></blockquote><p>The tradeoff is speed: starting a container takes a few seconds. For this reason, Testcontainers tests belong in the integration tier, not the unit tier.</p><h5><strong>Test naming conventions and organization</strong></h5><p>How you name and organize tests matters more than you might think. In six months, when a test fails in CI, the test name is the first thing you will see. A good name tells you exactly what broke without reading the code.</p><p>Here are some conventions that work well:</p><p><strong>File organization:</strong></p><pre><code>src/
  utils.ts
  utils.test.ts        # Co-located with the source file
  app.ts
  app.test.ts
  weather.ts
  weather.test.ts</code></pre><p>Co-locating tests with source files makes it obvious which file a test covers. Some teams prefer a separate <code>__tests__</code> directory, but co-location has the advantage that when you rename or move a file, the test moves with it.</p><p><strong>Naming patterns:</strong></p><pre><code>// Good: Describes the behavior clearly
describe("slugify", () =&gt; {
  it("converts spaces to dashes", () =&gt; {});
  it("removes special characters", () =&gt; {});
  it("handles empty string", () =&gt; {});
});

// Bad: Vague or implementation-focused
describe("slugify", () =&gt; {
  it("works", () =&gt; {});
  it("test 1", () =&gt; {});
  it("uses regex", () =&gt; {}); // who cares about the implementation?
});</code></pre><p>The test name should answer: &#8220;What behavior does this test verify?&#8221; When it fails, the output should read like a bug report: <code>slugify &gt; removes special characters: FAILED</code>.</p><h5><strong>What coverage actually means</strong></h5><p>Code coverage measures what percentage of your code is executed when your tests run. You can generate a coverage report with Vitest:</p><pre><code>npx vitest run --coverage</code></pre><p>This gives you metrics like:</p><blockquote><ul><li><p><strong>Line coverage</strong>: What percentage of lines were executed?</p></li><li><p><strong>Branch coverage</strong>: What percentage of if/else paths were taken?</p></li><li><p><strong>Function coverage</strong>: What percentage of functions were called?</p></li><li><p><strong>Statement coverage</strong>: What percentage of statements were executed?</p></li></ul></blockquote><p>A coverage report might look like this:</p><pre><code># ------------------|---------|----------|---------|---------|
# File              | % Stmts | % Branch | % Funcs | % Lines |
# ------------------|---------|----------|---------|---------|
# src/utils.ts      |   100   |   100    |   100   |   100   |
# src/weather.ts    |    85   |    75    |   100   |    85   |
# src/app.ts        |    92   |    80    |   100   |    92   |
# ------------------|---------|----------|---------|---------|</code></pre><p><strong>Why 100% coverage is a trap:</strong></p><p>Coverage tells you what code was executed, not what code was tested correctly. Consider this:</p><pre><code>// This test has 100% coverage of the add function
function add(a: number, b: number): number {
  return a + b;
}

it("covers the add function", () =&gt; {
  add(1, 2); // Look, we called it! 100% coverage!
  // But we never checked the result...
});</code></pre><p>That test executes every line of <code>add</code> but proves nothing. The function could return <code>"banana"</code> and the test would still pass. Coverage without meaningful assertions is theater.</p><p><strong>What metrics to watch instead:</strong></p><blockquote><ul><li><p><strong>Mutation testing</strong>: Tools like Stryker modify your code (change <code>+</code> to <code>-</code>, remove conditionals) and check if any tests fail. If a mutation survives, your tests have a blind spot. This is far more meaningful than line coverage.</p></li><li><p><strong>Branch coverage over line coverage</strong>: Branch coverage catches untested conditional paths. A function with an if/else might have 100% line coverage but only 50% branch coverage if you never test the else path.</p></li><li><p><strong>Test failure rate in CI</strong>: If tests never fail, they might not be testing anything meaningful. If they fail constantly, they might be flaky. A healthy test suite fails occasionally when real bugs are introduced.</p></li><li><p><strong>Time to detection</strong>: How quickly do tests catch a real bug after it is introduced? This is the metric that actually matters for DevOps.</p></li></ul></blockquote><p>A reasonable coverage target is somewhere between 70% and 90%. Anything above 90% usually means you are writing tests for trivial code just to hit a number.</p><h5><strong>When to NOT write tests</strong></h5><p>Testing everything is not the goal. Testing the right things is. Here are cases where writing tests adds cost without meaningful value:</p><blockquote><ul><li><p><strong>Generated code</strong>: If a tool generates your API client, ORM models, or GraphQL types, do not test the generation output. Test the code that uses them.</p></li><li><p><strong>Simple getters and setters</strong>: A function that just returns a property does not need a test. If you feel the need to test it, the function is probably too simple to break.</p></li><li><p><strong>Framework internals</strong>: Do not test that Express routes requests or that React renders components. Those are the framework&#8217;s job. Test your logic that runs inside the framework.</p></li><li><p><strong>Third-party libraries</strong>: Do not test that <code>lodash.groupBy</code> works correctly. The library maintainers already did that.</p></li><li><p><strong>Configuration files</strong>: JSON configs, environment variable listings, and static data do not need unit tests.</p></li></ul></blockquote><p>Focus your testing effort where bugs are most likely and most expensive: business logic, data transformations, edge cases in parsing, and integration points between systems.</p><h5><strong>Running tests in CI</strong></h5><p>We will cover CI/CD in detail in the next article, but here is a preview of what running tests in GitHub Actions looks like:</p><pre><code># .github/workflows/test.yml
name: Tests

on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

jobs:
  test:
    runs-on: ubuntu-latest

    steps:
      - uses: actions/checkout@v4

      - uses: actions/setup-node@v4
        with:
          node-version: "22"
          cache: "npm"

      - run: npm ci

      - run: npm test

      - run: npm run test:coverage

      - name: Upload coverage report
        uses: actions/upload-artifact@v4
        with:
          name: coverage-report
          path: coverage/</code></pre><p>This workflow runs on every push and pull request. If any test fails, the CI run fails and the PR cannot be merged (assuming you have branch protection enabled). This is the gate we talked about earlier: code does not reach production unless it passes the tests.</p><p>Key things to notice:</p><blockquote><ul><li><p><strong><code>npm ci</code></strong> instead of <code>npm install</code>: this installs exact versions from <code>package-lock.json</code>, ensuring reproducible builds.</p></li><li><p><strong>Separate test and coverage steps</strong>: run tests first for fast feedback, then coverage as a separate step.</p></li><li><p><strong>Upload artifacts</strong>: coverage reports are saved so you can download and review them later.</p></li></ul></blockquote><p>We will expand on this significantly in article five, covering caching, matrix builds, parallel test execution, and more.</p><h5><strong>Putting it all together</strong></h5><p>Let&#8217;s review what a well-tested project looks like. Here is the full directory structure:</p><pre><code>testing-demo/
  package.json
  tsconfig.json
  src/
    utils.ts              # Pure utility functions
    utils.test.ts         # Unit tests for utils
    weather.ts            # Module with external dependency
    weather.test.ts       # Unit tests with mocks
    app.ts                # Express application
    app.test.ts           # Integration tests with Supertest</code></pre><p>Each layer of the pyramid is covered:</p><blockquote><ul><li><p><strong>Unit tests</strong> (<code>utils.test.ts</code>, <code>weather.test.ts</code>): Fast, isolated, no external dependencies. These catch logic bugs.</p></li><li><p><strong>Integration tests</strong> (<code>app.test.ts</code>): Test the HTTP layer end to end (within the app). These catch wiring bugs.</p></li><li><p><strong>E2E tests</strong> (not shown here): Would use a tool like Playwright or Cypress to test the full stack through a browser.</p></li></ul></blockquote><p>The testing workflow in a DevOps pipeline looks like this:</p><blockquote><ol><li><p>Developer pushes code.</p></li><li><p>CI runs unit tests (seconds).</p></li><li><p>CI runs integration tests (seconds to minutes).</p></li><li><p>CI runs E2E tests (minutes).</p></li><li><p>If all pass, the code is eligible for deployment.</p></li><li><p>If any fail, the pipeline stops and the developer is notified.</p></li></ol></blockquote><p>This is the fast feedback loop that makes DevOps work. You find bugs in minutes, not days.</p><h5><strong>Closing notes</strong></h5><p>Testing is not optional in a DevOps workflow. It is the foundation that makes everything else possible: continuous integration, continuous deployment, and the confidence to ship changes multiple times a day.</p><p>Start with unit tests. They are the cheapest and give you the most value per line of test code. Add integration tests for your API endpoints and critical data flows. Use E2E tests sparingly for your most important user journeys.</p><p>Do not chase coverage numbers. Focus on testing behavior that matters: business logic, edge cases, and integration points. A well-placed test that catches a real bug is worth more than a hundred tests that just inflate a coverage metric.</p><p>In the next article, we will take these tests and wire them into a proper CI/CD pipeline with GitHub Actions. You will see how to run tests automatically, cache dependencies for speed, and set up branch protection so untested code never reaches production.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Version Control for Teams]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-version-control-for-teams</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-version-control-for-teams</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Mon, 27 Apr 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Qp5f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Qp5f!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Qp5f!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!Qp5f!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!Qp5f!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!Qp5f!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Qp5f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043148?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Qp5f!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!Qp5f!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!Qp5f!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!Qp5f!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30e50e81-3ab2-4b54-a944-5bb1590f1615_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article 3 of the DevOps from Zero to Hero series. Now it is time to talk about one of the most critical tools in any engineering team: version control for teams.</p><p>If you are working alone, you can get away with committing directly to main. But the moment a second person touches the same codebase, you need structure, conventions, and guardrails to avoid stepping on each other&#8217;s toes. In this article we will cover branching strategies, pull requests, conventional commits, protected branches, and merge strategies.</p><h5><strong>Why Version Control Matters in DevOps</strong></h5><p>Version control is the foundation for everything you will build in DevOps:</p><blockquote><ul><li><p><strong>Collaboration</strong> Multiple people can work on the same codebase simultaneously without overwriting each other&#8217;s changes</p></li><li><p><strong>Auditability</strong> Every change is recorded with a timestamp, author, and message. You can trace exactly what changed and who changed it</p></li><li><p><strong>Rollback</strong> Made a bad deployment? Revert to a previous known-good state in seconds</p></li><li><p><strong>Automation</strong> CI/CD pipelines trigger based on Git events. Without version control, you cannot automate builds, tests, or deployments</p></li><li><p><strong>Code review</strong> Pull requests create a structured process for reviewing changes before they reach production</p></li></ul></blockquote><h5><strong>Branching Strategies</strong></h5><p>A branching strategy defines how your team uses branches to develop, test, and release software. Let&#8217;s look at the three most common ones.</p><p><strong>Trunk-Based Development</strong></p><p>Everyone commits to a single branch (<code>main</code>). Feature branches are short-lived, lasting no more than a day or two.</p><pre><code>main &#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;
            \       /   \     /
feature-a    &#9679;&#9472;&#9472;&#9472;&#9679;     feature-b</code></pre><blockquote><ul><li><p><strong>Short-lived branches</strong> Features are broken into small pieces merged within a day or two</p></li><li><p><strong>Feature flags</strong> Incomplete features are hidden behind flags so they can be merged safely</p></li><li><p><strong>Continuous integration</strong> Everyone integrates frequently, reducing merge conflicts</p></li></ul></blockquote><p>This is the strategy I recommend for most teams. Companies like Google and Netflix use it at scale.</p><p><strong>GitFlow</strong></p><p>Uses multiple long-lived branches: <code>main</code>, <code>develop</code>, <code>feature/*</code>, <code>release/*</code>, and <code>hotfix/*</code>. Designed for teams with scheduled releases.</p><pre><code>main     &#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;
               \                 /
develop   &#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9679;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;&#9472;
              \   /       \       /
feature-a      &#9679;&#9472;&#9472;&#9679;        release/1.0</code></pre><p>Good for versioned software (libraries, mobile apps), but adds unnecessary complexity for web applications deployed continuously.</p><p><strong>GitHub Flow</strong></p><p>Simplified workflow: <code>main</code> and short-lived feature branches. Open a PR, get a review, merge. That is it.</p><blockquote><ul><li><p><strong>Simple</strong> Only two branch types: <code>main</code> and feature branches</p></li><li><p><strong>PR driven</strong> Every change goes through a pull request</p></li><li><p><strong>Deploy from main</strong> The main branch is always deployable</p></li></ul></blockquote><p>For most teams building web applications, use trunk-based development or GitHub Flow. Use GitFlow only if you genuinely need structured release cycles.</p><h5><strong>Pull Requests: The Heart of Team Collaboration</strong></h5><p>A pull request is more than a merge request. It is a conversation about the changes you are proposing. Good PRs are the single most important practice for maintaining code quality.</p><p><strong>What a Good PR Looks Like</strong></p><pre><code>## What
Brief description of what this PR does.

## Why
Why is this change needed? Link to the issue or ticket.

## How
High-level overview of the approach.

## Testing
How was this tested? Include relevant test commands or screenshots.

## Checklist
- [ ] Tests pass locally
- [ ] Documentation updated if needed
- [ ] No breaking changes (or migration path documented)</code></pre><p><strong>Review Best Practices</strong></p><blockquote><ul><li><p><strong>Be constructive</strong> Suggest improvements, offer alternatives when you disagree</p></li><li><p><strong>Focus on what matters</strong> Architecture and correctness over style preferences. Let the linter handle formatting</p></li><li><p><strong>Ask questions</strong> &#8220;Could you explain why this approach was chosen?&#8221; beats &#8220;This is wrong&#8221;</p></li><li><p><strong>Review promptly</strong> Blocked PRs kill momentum. Review within hours, not days</p></li><li><p><strong>Keep PRs small</strong> Aim for under 400 lines. Large PRs get rubber-stamped</p></li></ul></blockquote><h5><strong>Conventional Commits</strong></h5><p>Conventional commits are a specification for structured commit messages that enable automated changelogs and version bumps.</p><p><strong>The Format</strong></p><pre><code>&lt;type&gt;(&lt;scope&gt;): &lt;description&gt;

[optional body]
[optional footer(s)]</code></pre><p>Common types: <code>feat</code>, <code>fix</code>, <code>docs</code>, <code>style</code>, <code>refactor</code>, <code>test</code>, <code>chore</code>, <code>ci</code>, <code>perf</code>.</p><p><strong>Examples</strong></p><pre><code>feat(auth): add OAuth2 login with GitHub
fix(api): handle null response from payment gateway
docs(readme): update installation instructions for v2
chore(deps): bump phoenix_live_view from 1.0.0 to 1.1.0
refactor(database): extract connection pooling into separate module
ci(github-actions): add Elixir formatter check to PR workflow
feat(notifications)!: redesign notification system

BREAKING CHANGE: notification payloads now use snake_case keys</code></pre><blockquote><ul><li><p><strong>Automated changelogs</strong> Tools like <code>release-please</code> generate changelogs from commit history</p></li><li><p><strong>Semantic versioning</strong> Commit type determines patch (fix), minor (feat), or major (breaking change)</p></li><li><p><strong>Readable history</strong> <code>git log --oneline</code> becomes a clear story of what happened</p></li></ul></blockquote><h5><strong>Protected Branches</strong></h5><p>Protected branches prevent dangerous actions on important branches like <code>main</code>. Here is how to set them up on GitHub:</p><ol><li><p>Go to your repository, then <strong>Settings</strong> &gt; <strong>Branches</strong></p></li><li><p>Click <strong>Add branch protection rule</strong>, set pattern to <code>main</code></p></li><li><p>Configure the rules:</p></li></ol><pre><code>[x] Require a pull request before merging
    [x] Require approvals: 1
    [x] Dismiss stale approvals when new commits are pushed

[x] Require status checks to pass before merging
    [x] Require branches to be up to date before merging

[x] Do not allow force pushes
[x] Do not allow deletions
[x] Do not allow bypassing the above settings</code></pre><p>You can also do this with the GitHub CLI:</p><pre><code>gh api repos/{owner}/{repo}/branches/main/protection \
  --method PUT \
  --input - &lt;&lt;'EOF'
{
  "required_status_checks": {
    "strict": true,
    "contexts": ["ci/tests", "ci/lint"]
  },
  "enforce_admins": true,
  "required_pull_request_reviews": {
    "dismiss_stale_reviews": true,
    "required_approving_review_count": 1
  },
  "restrictions": null,
  "allow_force_pushes": false,
  "allow_deletions": false
}
EOF</code></pre><p><strong>CODEOWNERS File</strong></p><p>Define who reviews changes to specific parts of the codebase:</p><pre><code># .github/CODEOWNERS
* @your-team/backend
/assets/ @your-team/frontend
/terraform/ @your-team/platform
/.github/workflows/ @your-team/platform @your-team/leads</code></pre><h5><strong>Merge Strategies</strong></h5><p>When merging a PR on GitHub, you have three options:</p><p><strong>Merge Commit</strong> preserves all individual commits plus a merge commit. Full history, but can get messy.</p><pre><code># * abc1234 Merge pull request #42
# |\
# | * def5678 fix: handle edge case
# | * ghi9012 feat: add JWT validation
# |/
# * mno7890 Previous commit</code></pre><p><strong>Squash and Merge</strong> combines all commits into one on main. Clean, linear history.</p><pre><code># * abc1234 feat(auth): add JWT authentication (#42)
# * mno7890 Previous commit</code></pre><p><strong>Rebase and Merge</strong> replays individual commits on top of main. Linear history with full commit detail.</p><pre><code># * abc1234 fix: handle edge case
# * def5678 feat: add JWT validation
# * ghi9012 feat: add login endpoint
# * mno7890 Previous commit</code></pre><p>For most teams, <strong>squash and merge</strong> is the sweet spot. It keeps main clean and encourages small, focused PRs.</p><h5><strong>Practical Tips</strong></h5><p><strong>Meaningful branch names</strong> with consistent prefixes:</p><pre><code>feature/add-oauth-login
fix/null-pointer-in-payment
chore/upgrade-elixir-to-1.17
docs/update-api-reference</code></pre><p><strong>Atomic commits</strong> where each commit represents one logical change:</p><pre><code>git add lib/auth/session.ex
git commit -m "fix(auth): prevent session fixation on password reset"

git add test/auth/session_test.exs
git commit -m "test(auth): add regression test for session fixation"</code></pre><p><strong>Rebase before merging</strong> to keep your branch up to date:</p><pre><code>git fetch origin
git rebase origin/main
# Use --force-with-lease instead of --force (safer)
git push --force-with-lease origin feature/add-auth</code></pre><p><strong>Clean up merged branches</strong> to avoid clutter:</p><pre><code>git fetch --prune
git branch --merged main | grep -v "main" | xargs git branch -d</code></pre><h5><strong>Closing notes</strong></h5><p>Version control is the backbone of how teams collaborate and how code reaches production safely. Getting your Git workflow right early saves enormous amounts of pain down the road.</p><p>The key takeaways: use trunk-based development or GitHub Flow, keep PRs small, adopt conventional commits, protect your main branch, and use squash merges. Start simple, add complexity only when you need it.</p><p>In the next article, we will dive into CI/CD pipelines. Stay tuned!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and examples in the <a href="https://github.com/kainlite/tr">repository here</a>.</p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: Your First TypeScript API with Express and Docker]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-your-first-typescript-api</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-your-first-typescript-api</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Fri, 24 Apr 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4tmN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4tmN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4tmN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!4tmN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!4tmN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!4tmN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4tmN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043150?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4tmN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!4tmN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!4tmN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!4tmN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F565e0fc1-2da8-40cb-944d-9afdb902e533_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>Welcome to article two of the DevOps from Zero to Hero series. In the first article we set up our development environment and got familiar with the basic tools. Now it is time to build something real: a REST API that we can deploy, test, and iterate on throughout the rest of the series.</p><p>We are going to build a simple task tracker API using TypeScript and Express. Nothing fancy, just CRUD operations on an in-memory array. The goal is not to build a production-grade app right now, but to have a working API that we can containerize, deploy, and improve in future articles.</p><p>After the API is working, we will write a Dockerfile using multi-stage builds, set up a <code>.dockerignore</code>, run the container as a non-root user, add a health check endpoint, and wire everything up with Docker Compose for local development with hot reload.</p><p>Let&#8217;s get into it.</p><h5><strong>Why TypeScript and Express?</strong></h5><p>You might wonder why we are not using Python, Go, or something else. TypeScript with Express is one of the most common stacks you will encounter in the wild. It has a massive ecosystem, the tooling is mature, and the concepts translate directly to other languages and frameworks.</p><p>For DevOps, the language itself matters less than understanding how to build, test, package, and deploy applications. We picked TypeScript because it gives us type safety without too much ceremony, and Express because it is minimal enough that we can focus on the DevOps side of things.</p><h5><strong>Project setup</strong></h5><p>First, create a new directory and initialize the project:</p><pre><code>mkdir task-api &amp;&amp; cd task-api
npm init -y</code></pre><p>Install the dependencies we need:</p><pre><code>npm install express
npm install -D typescript @types/express @types/node ts-node nodemon</code></pre><p>Now create the TypeScript configuration. This tells the compiler how to process our code:</p><pre><code>// tsconfig.json
{
  "compilerOptions": {
    "target": "ES2020",
    "module": "commonjs",
    "lib": ["ES2020"],
    "outDir": "./dist",
    "rootDir": "./src",
    "strict": true,
    "esModuleInterop": true,
    "skipLibCheck": true,
    "forceConsistentCasingInFileNames": true,
    "resolveJsonModule": true,
    "declaration": true,
    "declarationMap": true,
    "sourceMap": true
  },
  "include": ["src/**/*"],
  "exclude": ["node_modules", "dist"]
}</code></pre><p>Update your <code>package.json</code> scripts section:</p><pre><code>{
  "scripts": {
    "build": "tsc",
    "start": "node dist/index.js",
    "dev": "nodemon --watch src --ext ts --exec ts-node src/index.ts"
  }
}</code></pre><p>Create the source directory:</p><pre><code>mkdir src</code></pre><h5><strong>Defining the task model</strong></h5><p>Let&#8217;s start with a simple type definition for our tasks. Create <code>src/types.ts</code>:</p><pre><code>// src/types.ts
export interface Task {
  id: number;
  title: string;
  description: string;
  completed: boolean;
  createdAt: string;
  updatedAt: string;
}

export interface CreateTaskRequest {
  title: string;
  description?: string;
}

export interface UpdateTaskRequest {
  title?: string;
  description?: string;
  completed?: boolean;
}</code></pre><p>This gives us a clear contract for what a task looks like and what data we expect when creating or updating one.</p><h5><strong>Building the API</strong></h5><p>Now let&#8217;s build the actual API. Create <code>src/index.ts</code>:</p><pre><code>// src/index.ts
import express, { Request, Response } from "express";
import { Task, CreateTaskRequest, UpdateTaskRequest } from "./types";

const app = express();
const PORT = process.env.PORT || 3000;

// Middleware
app.use(express.json());

// In-memory storage
let tasks: Task[] = [];
let nextId = 1;

// Health check endpoint
app.get("/health", (_req: Request, res: Response) =&gt; {
  res.json({
    status: "healthy",
    uptime: process.uptime(),
    timestamp: new Date().toISOString(),
  });
});

// GET /tasks - List all tasks
app.get("/tasks", (_req: Request, res: Response) =&gt; {
  res.json({
    data: tasks,
    count: tasks.length,
  });
});

// GET /tasks/:id - Get a single task
app.get("/tasks/:id", (req: Request, res: Response) =&gt; {
  const task = tasks.find((t) =&gt; t.id === parseInt(req.params.id));
  if (!task) {
    res.status(404).json({ error: "Task not found" });
    return;
  }
  res.json({ data: task });
});

// POST /tasks - Create a new task
app.post("/tasks", (req: Request, res: Response) =&gt; {
  const body: CreateTaskRequest = req.body;

  if (!body.title || body.title.trim() === "") {
    res.status(400).json({ error: "Title is required" });
    return;
  }

  const now = new Date().toISOString();
  const task: Task = {
    id: nextId++,
    title: body.title.trim(),
    description: body.description?.trim() || "",
    completed: false,
    createdAt: now,
    updatedAt: now,
  };

  tasks.push(task);
  res.status(201).json({ data: task });
});

// PUT /tasks/:id - Update a task
app.put("/tasks/:id", (req: Request, res: Response) =&gt; {
  const taskIndex = tasks.findIndex((t) =&gt; t.id === parseInt(req.params.id));
  if (taskIndex === -1) {
    res.status(404).json({ error: "Task not found" });
    return;
  }

  const body: UpdateTaskRequest = req.body;
  const existing = tasks[taskIndex];

  const updated: Task = {
    ...existing,
    title: body.title?.trim() ?? existing.title,
    description: body.description?.trim() ?? existing.description,
    completed: body.completed ?? existing.completed,
    updatedAt: new Date().toISOString(),
  };

  tasks[taskIndex] = updated;
  res.json({ data: updated });
});

// DELETE /tasks/:id - Delete a task
app.delete("/tasks/:id", (req: Request, res: Response) =&gt; {
  const taskIndex = tasks.findIndex((t) =&gt; t.id === parseInt(req.params.id));
  if (taskIndex === -1) {
    res.status(404).json({ error: "Task not found" });
    return;
  }

  const deleted = tasks.splice(taskIndex, 1)[0];
  res.json({ data: deleted, message: "Task deleted" });
});

// Start the server
app.listen(PORT, () =&gt; {
  console.log(`Task API running on port ${PORT}`);
});

export default app;</code></pre><h5><strong>Testing the API locally</strong></h5><p>Start the development server:</p><pre><code>npm run dev</code></pre><p>You should see <code>Task API running on port 3000</code>. Now let&#8217;s test each endpoint with curl:</p><pre><code># Health check
curl http://localhost:3000/health | jq

# Create a task
curl -X POST http://localhost:3000/tasks \
  -H "Content-Type: application/json" \
  -d '{"title": "Learn Docker", "description": "Build and run containers"}' | jq

# Create another task
curl -X POST http://localhost:3000/tasks \
  -H "Content-Type: application/json" \
  -d '{"title": "Write Dockerfile", "description": "Multi-stage build"}' | jq

# List all tasks
curl http://localhost:3000/tasks | jq

# Get a single task
curl http://localhost:3000/tasks/1 | jq

# Update a task
curl -X PUT http://localhost:3000/tasks/1 \
  -H "Content-Type: application/json" \
  -d '{"completed": true}' | jq

# Delete a task
curl -X DELETE http://localhost:3000/tasks/2 | jq</code></pre><p>You should see proper JSON responses for each request. The health check returns the server uptime, the POST returns the created task with an auto-incremented ID, and so on.</p><h5><strong>The Dockerfile</strong></h5><p>Now we get to the fun part. We are going to containerize this API using Docker best practices.</p><p>First, let&#8217;s talk about why multi-stage builds matter. A typical TypeScript project has development dependencies (the compiler, type definitions, nodemon) that we do not need at runtime. With multi-stage builds, we compile in one stage and copy only the output to a smaller final image. This means smaller images, faster pulls, and a smaller attack surface.</p><p>Create the Dockerfile:</p><pre><code># Stage 1: Build
FROM node:20-alpine AS builder

WORKDIR /app

# Copy package files first for better layer caching
COPY package*.json ./

# Install all dependencies (including devDependencies for building)
RUN npm ci

# Copy source code
COPY tsconfig.json ./
COPY src ./src

# Compile TypeScript
RUN npm run build

# Stage 2: Production
FROM node:20-alpine AS production

# Add a non-root user
RUN addgroup -g 1001 appgroup &amp;&amp; \
    adduser -u 1001 -G appgroup -s /bin/sh -D appuser

WORKDIR /app

# Copy package files and install production-only dependencies
COPY package*.json ./
RUN npm ci --only=production &amp;&amp; npm cache clean --force

# Copy compiled output from builder stage
COPY --from=builder /app/dist ./dist

# Switch to non-root user
USER appuser

# Expose the port
EXPOSE 3000

# Set environment variable
ENV NODE_ENV=production

# Health check using the /health endpoint
HEALTHCHECK --interval=30s --timeout=3s --start-period=5s --retries=3 \
  CMD wget --no-verbose --tries=1 --spider http://localhost:3000/health || exit 1

# Start the application
CMD ["node", "dist/index.js"]</code></pre><p>Let&#8217;s break down what each part does:</p><blockquote><ul><li><p><strong>Multi-stage build</strong> We use two stages. The first installs all dependencies and compiles TypeScript. The second only has production dependencies and the compiled JavaScript. This keeps the final image small.</p></li><li><p><strong>Alpine base</strong> We use <code>node:20-alpine</code> instead of <code>node:20</code>. Alpine is a minimal Linux distribution that produces much smaller images.</p></li><li><p><strong>Layer caching</strong> We copy <code>package*.json</code> before the source code. This means Docker can cache the <code>npm ci</code> layer and only reinstall dependencies when <code>package.json</code> changes.</p></li><li><p><strong>Non-root user</strong> Running as root inside a container is a security risk. We create a dedicated user and switch to it before starting the app.</p></li><li><p><strong>Health check</strong> Docker can monitor the container health by hitting our <code>/health</code> endpoint. Orchestrators like Kubernetes use this to know when to restart unhealthy containers.</p></li></ul></blockquote><h5><strong>The .dockerignore file</strong></h5><p>Just like <code>.gitignore</code> keeps files out of your repository, <code>.dockerignore</code> keeps files out of your Docker build context. This makes builds faster and prevents sensitive files from leaking into images.</p><p>Create <code>.dockerignore</code>:</p><pre><code>node_modules
dist
npm-debug.log
.git
.gitignore
.env
.env.*
*.md
.vscode
.idea
coverage
.nyc_output</code></pre><p>The most important entry is <code>node_modules</code>. Without this, Docker would copy your entire local <code>node_modules</code> directory into the build context, which is slow and unnecessary since we run <code>npm ci</code> inside the container anyway.</p><h5><strong>Building and running the container</strong></h5><p>Build the image:</p><pre><code>docker build -t task-api:latest .</code></pre><p>You should see Docker executing both stages. The first time takes a bit longer because it downloads the base image and installs dependencies. Subsequent builds are faster thanks to layer caching.</p><p>Check the image size:</p><pre><code>docker images task-api</code></pre><p>The Alpine-based multi-stage image should be around 130-150 MB. Compare that to a full <code>node:20</code> image which starts at over 900 MB before you even add your code.</p><p>Run the container:</p><pre><code>docker run -d --name task-api -p 3000:3000 task-api:latest</code></pre><p>Test it:</p><pre><code># Health check
curl http://localhost:3000/health | jq

# Create a task
curl -X POST http://localhost:3000/tasks \
  -H "Content-Type: application/json" \
  -d '{"title": "Running in Docker!"}' | jq</code></pre><p>Check the container health status:</p><pre><code>docker inspect --format='{{.State.Health.Status}}' task-api</code></pre><p>After about 30 seconds, it should show <code>healthy</code>.</p><p>Stop and remove the container when you are done:</p><pre><code>docker stop task-api &amp;&amp; docker rm task-api</code></pre><h5><strong>Docker Compose for local development</strong></h5><p>Running <code>docker build</code> and <code>docker run</code> every time you change code gets old fast. Docker Compose gives us a better workflow. We can define services, mount our source code as a volume, and get hot reload inside the container.</p><p>Create <code>docker-compose.yml</code>:</p><pre><code>services:
  api:
    build:
      context: .
      dockerfile: Dockerfile.dev
    ports:
      - "3000:3000"
    volumes:
      - ./src:/app/src
      - ./package.json:/app/package.json
    environment:
      - NODE_ENV=development
      - PORT=3000
    restart: unless-stopped</code></pre><p>We need a separate Dockerfile for development since we want <code>ts-node</code> and <code>nodemon</code> available. Create <code>Dockerfile.dev</code>:</p><pre><code>FROM node:20-alpine

WORKDIR /app

# Copy package files and install all dependencies
COPY package*.json ./
RUN npm ci

# Copy TypeScript config
COPY tsconfig.json ./

# Copy source code (will be overridden by volume mount)
COPY src ./src

# Expose the port
EXPOSE 3000

# Run with nodemon for hot reload
CMD ["npx", "nodemon", "--watch", "src", "--ext", "ts", "--exec", "ts-node", "src/index.ts"]</code></pre><p>Start the development environment:</p><pre><code>docker compose up</code></pre><p>Now edit <code>src/index.ts</code>, save the file, and watch nodemon restart automatically inside the container. Your changes appear without rebuilding the image. This is the development workflow you want: fast feedback loops while still running inside a container.</p><p>To run it in the background:</p><pre><code>docker compose up -d</code></pre><p>Check the logs:</p><pre><code>docker compose logs -f api</code></pre><p>Stop everything:</p><pre><code>docker compose down</code></pre><h5><strong>Why containers matter for DevOps</strong></h5><p>We just went from &#8220;code on my machine&#8221; to &#8220;code in a container.&#8221; This might seem like extra work for a simple API, but containers solve real problems that show up in every team:</p><blockquote><ul><li><p><strong>Reproducibility</strong> The container runs the same way on your laptop, in CI, and in production. No more &#8220;it works on my machine&#8221; conversations.</p></li><li><p><strong>Consistency</strong> Everyone on the team uses the same Node.js version, the same OS, the same dependencies. The Dockerfile is the single source of truth.</p></li><li><p><strong>Isolation</strong> Your app runs in its own filesystem and network namespace. It does not conflict with other services on the same machine.</p></li><li><p><strong>Portability</strong> The image runs anywhere Docker runs: local machines, cloud VMs, Kubernetes clusters. You build once and deploy anywhere.</p></li><li><p><strong>Immutability</strong> Once built, the image does not change. You do not SSH into production and tweak files. You build a new image and deploy it.</p></li></ul></blockquote><p>These properties are the foundation of modern DevOps. Every tool and practice we cover in this series builds on top of containers. CI/CD pipelines build container images. Kubernetes orchestrates them. GitOps tracks which image version runs where. Without containers, none of that works as smoothly.</p><h5><strong>Project structure recap</strong></h5><p>At this point, your project should look like this:</p><pre><code>task-api/
&#9500;&#9472;&#9472; src/
&#9474;   &#9500;&#9472;&#9472; index.ts
&#9474;   &#9492;&#9472;&#9472; types.ts
&#9500;&#9472;&#9472; .dockerignore
&#9500;&#9472;&#9472; docker-compose.yml
&#9500;&#9472;&#9472; Dockerfile
&#9500;&#9472;&#9472; Dockerfile.dev
&#9500;&#9472;&#9472; package.json
&#9500;&#9472;&#9472; package-lock.json
&#9492;&#9472;&#9472; tsconfig.json</code></pre><h5><strong>Closing notes</strong></h5><p>In this article we built a complete REST API with TypeScript and Express, then containerized it using Docker best practices. We covered multi-stage builds, non-root users, health checks, <code>.dockerignore</code>, and Docker Compose for local development.</p><p>The API itself is intentionally simple. It stores tasks in memory, which means all data disappears when the container restarts. That is fine for now. In a future article we will add a real database and learn how to manage data persistence with containers.</p><p>In the next article, we will set up a CI/CD pipeline that automatically builds our Docker image, runs tests, and pushes the image to a container registry. That is where the DevOps workflow really starts to come together.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item><item><title><![CDATA[DevOps from Zero to Hero: What It Actually Means and Why You Should Care]]></title><description><![CDATA[Introduction]]></description><link>https://segfaultpw.substack.com/p/devops-from-zero-to-hero-what-it-actually-means</link><guid isPermaLink="false">https://segfaultpw.substack.com/p/devops-from-zero-to-hero-what-it-actually-means</guid><dc:creator><![CDATA[Gabriel]]></dc:creator><pubDate>Tue, 21 Apr 2026 00:00:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SWdQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SWdQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SWdQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!SWdQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!SWdQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!SWdQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SWdQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp" width="1456" height="761" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:761,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115728,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://segfaultpw.substack.com/i/201043152?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SWdQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 424w, https://substackcdn.com/image/fetch/$s_!SWdQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 848w, https://substackcdn.com/image/fetch/$s_!SWdQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 1272w, https://substackcdn.com/image/fetch/$s_!SWdQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F571b23a1-c0ca-49e9-ae0d-4080a5258a33_3240x1693.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h5><strong>Introduction</strong></h5><p>This is the first article in a twenty-part series called &#8220;DevOps from Zero to Hero.&#8221; The goal is to take you from knowing nothing about DevOps to being comfortable with the tools and practices that modern teams use every day. We will use TypeScript, AWS, Kubernetes, and GitHub Actions throughout the series, building real things along the way.</p><p>But before we touch any tools, we need to understand what DevOps actually is. This word gets thrown around a lot. Job postings ask for &#8220;DevOps Engineers,&#8221; companies buy &#8220;DevOps tools,&#8221; and somehow everyone has a different definition. In this article we are going to cut through the noise and talk about what DevOps really means, where it came from, how to measure it, and what it is definitely not.</p><p>Let&#8217;s get into it.</p><h5><strong>What is DevOps?</strong></h5><p>DevOps is not a tool. It is not a job title. It is not a team you create so developers can stop caring about production. DevOps is a combination of cultural practices, processes, and tools that increases an organization&#8217;s ability to deliver software faster and more reliably.</p><p>The simplest way to think about it: DevOps is about removing the walls between the people who write code and the people who run it in production.</p><p>There are three pillars to DevOps:</p><blockquote><ul><li><p><strong>Culture</strong>: Teams share responsibility for the full lifecycle of their software, from writing it to running it</p></li><li><p><strong>Practices</strong>: Continuous integration, continuous delivery, infrastructure as code, monitoring, and fast feedback loops</p></li><li><p><strong>Tools</strong>: The automation that makes those practices possible at scale</p></li></ul></blockquote><p>If you only adopt the tools without changing how your teams work, you are not doing DevOps. You are just automating the same broken process. This is a critical point that many organizations miss.</p><h5><strong>A brief history: the wall of confusion</strong></h5><p>To understand why DevOps exists, you need to know what came before it. For decades, software organizations had two separate groups:</p><blockquote><ul><li><p><strong>Development (Dev)</strong>: Writes the code, ships features, moves fast, wants to deploy often</p></li><li><p><strong>Operations (Ops)</strong>: Runs the servers, keeps things stable, moves carefully, wants to deploy never</p></li></ul></blockquote><p>These two groups had completely different incentives. Dev wanted change because change meant new features. Ops wanted stability because change meant risk. The handoff between them was called &#8220;the wall of confusion.&#8221; Dev would throw code over the wall, Ops would try to figure out how to run it, and when things broke, everyone blamed each other.</p><p>This created a painful cycle:</p><blockquote><ul><li><p>Deployments were rare (monthly or quarterly) because they were risky and stressful</p></li><li><p>Each deployment was huge because all the changes piled up</p></li><li><p>Huge deployments meant more things could go wrong</p></li><li><p>When things went wrong, it took forever to figure out which change caused the problem</p></li><li><p>So deployments became even more rare, and the cycle continued</p></li></ul></blockquote><p>In 2008 and 2009, a few people started talking about breaking this cycle. Patrick Debois organized the first &#8220;DevOpsDays&#8221; conference in Ghent, Belgium in 2009. The idea was simple: what if Dev and Ops worked together instead of against each other? What if we deployed small changes frequently instead of big changes rarely? What if we automated everything that could be automated?</p><p>These ideas were not entirely new. Google had been practicing something similar internally for years (they later published it as Site Reliability Engineering). But the DevOps movement gave it a name and made it accessible to everyone, not just companies with Google-scale resources.</p><h5><strong>The DORA metrics: measuring DevOps performance</strong></h5><p>One of the most important contributions to the DevOps movement came from the DORA (DevOps Research and Assessment) team, led by Dr. Nicole Forsgren, Jez Humble, and Gene Kim. They spent years researching what separates high-performing teams from low-performing ones. Their findings were published in the book &#8220;Accelerate&#8221; and in annual State of DevOps reports.</p><p>They identified four key metrics that predict software delivery performance:</p><blockquote><ul><li><p><strong>Deployment Frequency</strong>: How often your team deploys to production. Elite teams deploy on demand, multiple times per day. Low performers deploy monthly or less.</p></li><li><p><strong>Lead Time for Changes</strong>: How long it takes from a code commit to that code running in production. Elite teams measure this in less than one hour. Low performers take between one and six months.</p></li><li><p><strong>Change Failure Rate</strong>: What percentage of deployments cause a failure in production that requires a fix (rollback, patch, etc.). Elite teams have a rate of 0-15%. Low performers hit 46-60%.</p></li><li><p><strong>Mean Time to Recovery (MTTR)</strong>: When something breaks in production, how long does it take to restore service? Elite teams recover in less than one hour. Low performers take between one week and one month.</p></li></ul></blockquote><p>Here is the key insight from their research: these four metrics are correlated. Teams that deploy more frequently also have lower failure rates and faster recovery times. Speed and stability are not enemies. They reinforce each other.</p><pre><code>Traditional thinking:
  "If we deploy more often, more things will break"

What DORA research actually shows:
  "Teams that deploy more often break fewer things AND recover faster"

Why? Because:
  - Smaller changes are easier to understand and debug
  - Frequent deployments mean faster feedback loops
  - Fast feedback loops mean problems get caught earlier
  - Earlier problems are cheaper and simpler to fix</code></pre><p>This might feel counterintuitive at first. But think about it this way: would you rather debug a deployment that contains 3 commits or one that contains 300? The answer is obvious. Deploying frequently forces you to keep changes small, and small changes are inherently less risky.</p><h5><strong>DevOps vs SRE vs Platform Engineering</strong></h5><p>You will hear these three terms used interchangeably, but they are distinct (and complementary) disciplines. Understanding how they relate will save you a lot of confusion.</p><p><strong>DevOps</strong> is the cultural movement. It is the philosophy that says Dev and Ops should work together, share responsibility, and use automation to deliver software faster and more reliably. DevOps is about principles: you own what you build, you automate everything you can, and you measure outcomes.</p><p><strong>Site Reliability Engineering (SRE)</strong> is one way to implement DevOps principles. Google created it in the early 2000s before the term &#8220;DevOps&#8221; even existed. SRE treats operations as a software engineering problem. SRE teams write code to automate operational work, define Service Level Objectives (SLOs) to measure reliability, and use error budgets to balance reliability with feature velocity.</p><p>Ben Treynor Sloss, the founder of Google&#8217;s SRE team, described it this way:</p><pre><code>"SRE is what happens when you ask a software engineer to design an operations function."</code></pre><p>If DevOps is the &#8220;what&#8221; (principles and culture), SRE is one answer to the &#8220;how&#8221; (specific practices and frameworks).</p><p><strong>Platform Engineering</strong> is the newest of the three. It emerged as organizations realized that asking every development team to fully own their infrastructure was not scaling. Platform Engineering teams build internal developer platforms (IDPs) that abstract away infrastructure complexity. Instead of every team learning Kubernetes, Terraform, and CI/CD pipelines from scratch, the platform team provides golden paths, templates, and self-service tools.</p><p>Think of it this way:</p><pre><code>DevOps says:     "You build it, you run it"
SRE says:        "Here are the practices and metrics to run it well"
Platform Eng:    "Here is a platform that makes running it easy"</code></pre><p>These three approaches are not competing. In a mature organization, they work together. DevOps provides the culture, SRE provides the reliability framework, and Platform Engineering provides the developer experience layer on top.</p><h5><strong>The DevOps toolchain</strong></h5><p>While DevOps is not just about tools, the tools do matter. They are what make the practices possible at scale. Here is the typical DevOps toolchain, organized by stage:</p><p><strong>Plan and track</strong></p><blockquote><ul><li><p>Issue trackers (GitHub Issues, Jira, Linear)</p></li><li><p>Project boards, documentation wikis</p></li></ul></blockquote><p><strong>Version control</strong></p><blockquote><ul><li><p>Git (GitHub, GitLab, Bitbucket)</p></li><li><p>Branching strategies, pull requests, code review</p></li></ul></blockquote><p><strong>Continuous Integration (CI)</strong></p><blockquote><ul><li><p>Automatically build, test, and validate every code change</p></li><li><p>Tools: GitHub Actions, GitLab CI, Jenkins, CircleCI</p></li></ul></blockquote><p><strong>Continuous Delivery/Deployment (CD)</strong></p><blockquote><ul><li><p>Automatically deploy validated changes to production</p></li><li><p>Tools: ArgoCD, Flux, Spinnaker, GitHub Actions</p></li></ul></blockquote><p><strong>Containers and orchestration</strong></p><blockquote><ul><li><p>Package applications consistently across environments</p></li><li><p>Tools: Docker, Kubernetes, ECS</p></li></ul></blockquote><p><strong>Infrastructure as Code (IaC)</strong></p><blockquote><ul><li><p>Define and manage infrastructure through code, not clicking in consoles</p></li><li><p>Tools: Terraform, Pulumi, AWS CDK, CloudFormation</p></li></ul></blockquote><p><strong>Monitoring and observability</strong></p><blockquote><ul><li><p>Know what is happening in production before your users tell you</p></li><li><p>Tools: Prometheus, Grafana, Datadog, OpenTelemetry</p></li></ul></blockquote><p><strong>Security</strong></p><blockquote><ul><li><p>Shift security left, automate scanning, manage secrets</p></li><li><p>Tools: Trivy, Snyk, HashiCorp Vault, GitHub security features</p></li></ul></blockquote><p>In this series we will focus on a specific subset of these tools: TypeScript for application code, GitHub Actions for CI/CD, Docker for containers, Kubernetes for orchestration, and AWS for cloud infrastructure. This stack is widely used, well documented, and gives you skills that transfer to almost any organization.</p><h5><strong>What this series will cover</strong></h5><p>Here is the roadmap for the twenty articles in this series:</p><blockquote><ul><li><p><strong>Article 1 (this one)</strong>: What DevOps actually means</p></li><li><p><strong>Articles 2-3</strong>: Version control with Git and GitHub workflows</p></li><li><p><strong>Articles 4-5</strong>: Containers with Docker, from basics to multi-stage builds</p></li><li><p><strong>Articles 6-8</strong>: CI/CD with GitHub Actions, from simple pipelines to advanced workflows</p></li><li><p><strong>Articles 9-11</strong>: Cloud fundamentals with AWS (networking, compute, storage)</p></li><li><p><strong>Articles 12-14</strong>: Kubernetes from scratch, deploying and managing real applications</p></li><li><p><strong>Articles 15-16</strong>: Infrastructure as Code with Terraform</p></li><li><p><strong>Articles 17-18</strong>: Monitoring, logging, and observability</p></li><li><p><strong>Article 19</strong>: Security practices and secrets management</p></li><li><p><strong>Article 20</strong>: Putting it all together, a complete DevOps pipeline from commit to production</p></li></ul></blockquote><p>Each article builds on the previous ones. By the end of the series, you will have built a complete pipeline that takes a TypeScript application from a git commit all the way to a production Kubernetes cluster on AWS, with automated testing, security scanning, monitoring, and alerting.</p><p><strong>Who is this for?</strong></p><blockquote><ul><li><p>Developers who want to understand what happens to their code after they push it</p></li><li><p>Junior engineers or students who want to learn modern DevOps practices from scratch</p></li><li><p>Ops people who want to adopt a more engineering-driven approach</p></li><li><p>Anyone who keeps hearing &#8220;DevOps&#8221; in meetings and wants to actually understand what it means</p></li></ul></blockquote><p>You do not need prior experience with any of the tools we will use. I will explain everything from the ground up. Basic programming knowledge and comfort with the command line are helpful but not strictly required.</p><h5><strong>What DevOps is NOT: common anti-patterns</strong></h5><p>Let&#8217;s close with something equally important: what DevOps is not. These are real anti-patterns that organizations fall into constantly.</p><p><strong>Anti-pattern 1: Renaming your Ops team to &#8220;DevOps&#8221;</strong></p><p>If you take your existing operations team, change their title to &#8220;DevOps Engineer,&#8221; and nothing else changes, you have not adopted DevOps. You have renamed a team. DevOps requires cultural change, not just a title change.</p><p><strong>Anti-pattern 2: Buying tools and calling it DevOps</strong></p><p>Purchasing a CI/CD platform, a container orchestrator, and a monitoring tool does not make you a DevOps organization. Tools without the right practices and culture are just expensive shelfware. I have seen organizations spend millions on tooling while their teams still deploy manually every two weeks.</p><p><strong>Anti-pattern 3: Creating a DevOps silo</strong></p><p>The irony of this one is painful. DevOps was created to break down silos between Dev and Ops. Some organizations responded by creating a third silo called &#8220;the DevOps team&#8221; that sits between Dev and Ops. Now you have three walls of confusion instead of one.</p><p><strong>Anti-pattern 4: All tools, no culture</strong></p><p>This is worth repeating because it is the most common mistake. If your developers write code and then throw it over the wall to someone else to deploy, you are not doing DevOps no matter what tools you use. DevOps means shared ownership. The team that builds the software is responsible for running it.</p><p><strong>Anti-pattern 5: DevOps means &#8220;developers do everything&#8221;</strong></p><p>DevOps does not mean firing your ops team and making developers manage servers. It means that development and operations work together, share knowledge, and both contribute to automation. Developers gain operational awareness, and ops engineers gain development skills. The goal is collaboration, not consolidation.</p><h5><strong>Closing notes</strong></h5><p>DevOps is, at its core, a simple idea: the people who build software and the people who run it should work together, share responsibility, and use automation to move faster without sacrificing stability. The DORA metrics prove that this approach works. Speed and reliability are not opposites. They go hand in hand.</p><p>In the next article, we will start getting practical. We will set up a development environment, create a TypeScript project, initialize a Git repository, and learn the version control fundamentals that everything else in this series will build on.</p><p>Hope you found this useful and enjoyed reading it, until next time!</p><h5><strong>Errata</strong></h5><p>If you spot any error or have any suggestion, please send me a message so it gets fixed.</p><p>Also, you can check the source code and changes in the <a href="https://github.com/kainlite/tr">sources here</a></p>]]></content:encoded></item></channel></rss>