Web Development

Stop Fighting CI/CD Logs Like a Caveman: Let Claude Do the Archaeology

A

Adil Sher

Author

Oct 3, 2026
5 min read
0 views
Stop Fighting CI/CD Logs Like a Caveman: Let Claude Do the Archaeology

Three months ago, I spent ninety minutes debugging a GitHub Actions failure that turned out to be a missing s3:GetObject permission in an IAM role. Ninety minutes. The actual fix was changing one line in a Terraform policy document. I scrolled through what felt like ten thousand lines of Docker layer caching, npm package installation spam, and environment variable dumps before finding the three-line error message buried near the end. I remember thinking: "There has to be a better way to do this."

Then I ran into a blog post about using Claude to automatically triage failed CI/CD pipelines, and something clicked. Not because I hadn't thought about using AI for this, I had. But because the author approached it with the right constraint: diagnose, don't deploy. This distinction matters more than most discussions about AI in DevOps acknowledge.

The Real Problem We're Solving

Let me be direct: CI/CD pipeline failures are asymmetric problems. Finding the root cause takes forever. Fixing it takes minutes. I've watched teams lose entire afternoons to log archaeology when the actual solution could be implemented in a Slack message. The pipeline runs fail, Slack pings explode, and someone gets stuck in the AWS console or GitHub Actions logs hunting for that one error signal among thousands of lines of noise.

The original article tackles this by building an agent that runs only when pipelines fail. When a GitHub Actions job explodes, a lightweight workflow grabs the logs, feeds them through Amazon Bedrock's Claude model, and posts a diagnosis directly to the pull request. No infrastructure to babysit. No persistent AWS keys sitting in repository secrets. Just temporary credentials, OIDC federation, and a Python script that knows how to read failure signals.

What I respect here is the restraint. The agent diagnoses and suggests fixes but never touches production. It's not auto-remediation, which would terrify me in most contexts. It's scaffolding for human decision-making.

The Technical Pattern That Actually Works

The author reveals something I've learned the hard way: raw logs destroy LLM reasoning. Feeding Claude twenty thousand lines of Docker build output and npm install logs wastes tokens and costs money while making the model less accurate. Instead, the system uses a distillation utility that strips ANSI color codes and hunts for error markers, traceback patterns, exit codes, IAM exceptions, Terraform failures.

When it finds a match, it extracts a window: fifteen lines before and thirty-five after the error. Suddenly you're down from fifty thousand tokens to maybe six hundred. Response times drop from fifteen seconds to two. The model stops hallucinating about warnings from line fifty and focuses on the actual failure.

# Simplified version of the distillation logic
def extract_failure_context(raw_log: str) -> str:
 patterns = [
 r"error[:\s]",
 r"traceback \(most recent call last\):",
 r"AccessDeniedException",
 r"failed with exit code \d+"
 ]
 
 lines = raw_log.splitlines()
 for idx, line in enumerate(lines):
 for pattern in patterns:
 if re.search(pattern, line, re.IGNORECASE):
 start = max(0, idx - 15)
 end = min(len(lines), idx + 35)
 return "\n".join(lines[start:end])

This is the kind of insight that comes from running this in production, not from a tutorial. You can't reason your way to this. You have to deploy it, watch it fail, and iterate.

My Take: Where I'd Push Back (Slightly)

I love the AWS OIDC federation approach, it's cleaner than static keys and I've been pushing teams toward it for eighteen months. But I wonder about the error pattern matching. In my infrastructure, I've got custom deployment tools that fail in ways Claude might not recognize. What happens when the distiller misses the actual failure? Does the agent send back a useless response that wastes the developer's time?

I'd probably add observability around the distiller itself. Log which patterns matched, how much context was extracted, and whether Claude's response was actually helpful. Then I'd integrate feedback so the error patterns evolve.

Also, and this is production paranoia, I want to see the drift over time. When new AWS services are launched or error message formats change, does this agent keep working? Or does it gradually become less useful as AWS evolves around it?

What's Next

I'm genuinely considering building something like this for our Islamabad team. We run Kubernetes clusters alongside AWS infrastructure, and our CI/CD logs are chaotic. But I want to start smaller: just ECS and Terraform failures. Then add observability around the diagnostic quality.

If you've built diagnostic tooling around your own pipelines, I want to hear what patterns you use. Are there error signatures that consistently trip up your team? That's where I'd focus first.

Source: This post was inspired by "Building an AWS DevOps Agent That Diagnoses but Never Deploys" by Dev.to. Read the original article

Written by Adil Sher

Full stack developer building high-traffic platforms, AI services, and custom web applications. Explore my portfolio, learn about my background, or get in touch.

Related Articles

I Wasted 3 Months Learning AI the Wrong Way, Here's What Actually Worked
Web Development Oct 1

I Wasted 3 Months Learning AI the Wrong Way, Here's What Actually Worked

I remember the exact moment I decided to "become an AI engineer." It was late 2023, everyone was talking about LLMs, and I had a successful web development background. How hard could it be? I spent the next three months jumping between Andrew Ng's machine learning course, LangCha...

I'm Done Betting On "Enterprise" Secrets Managers For Small Teams
Web Development Sep 30

I'm Done Betting On "Enterprise" Secrets Managers For Small Teams

Six months ago, I spent three days setting up HashiCorp Vault for a four-person team. We spent another week writing deployment scripts, managing encryption keys, and debugging authentication flows. By month two, nobody could remember how to rotate a key without breaking staging....