A little over four months ago I posted about Matty Code, an autonomous coding agent that we rolled out to allow Mattermost engineers to delegate tasks to an agent that would drive a task up to and through the finish line.
Over the last four months the landscape has changed quite a bit. Models have improved dramatically, tooling has added more features, more competition exists in the cloud agent space, and we’ve tweaked and improved the agent’s processes from internal feedback.
In the six months since it was initially rolled out, Matty (in some shape or form) has become a top-three contributor to the mattermost/mattermost repository, surpassing Claude, which is the local agent of choice for the vast majority of humans at Mattermost.
I wanted to take some time to write this post to reflect on Matty — what worked, what didn’t, what challenges we ran into (process, technical, and human) and how we overcame them, where I think the software development industry is headed, and where we hope to take Matty Code in the future.
The Human Side: Rollout, Trust, and Anxiety Over AI Slop
When I initially pitched Matty Code to the engineering team at Mattermost, it was met with heavy (and healthy) skepticism. Many engineers weren’t ready to completely let go, deferring a ticket in its entirety to an agent that would spit out a PR on the other end. Fears of reviewing AI slop PRs were rampant. No one wants to be a clean-up crew for a bot. Adding to this were concerns over ownership and accountability. If Matty caused a regression, who’s “fault” was it?
Even among the more AI-forward individuals in the org, a common question often came up: How is this any different than me just having my local Claude Code address the ticket?
What a lot of us forget is that even in the age of agents, there’s a real fixed cost (both literally and cognitively) to taking on a task yourself. You have to create the branch or worktree, get the instance running, repro the bug (or tell the agent to do it), send the agent in to fix it, review the fix, open the PR, address any findings from CodeRabbit, then send to another human for review.
While you’re waiting for review, you’re babysitting the PR through CI failures from flaky tests, merge conflicts, and human feedback (which cycles into CI failures and conflicts). All of this happens while you’re context switching into tasks that are actually important and need your full brain power and attention.
Matty eases both the literal and cognitive load by handling things end to end, freeing up time for you to think about the important things.
Evangelizing Matty was important to its success early on. I worked closely with some of the more AI-forward engineers at the company to encourage them to delegate more and more tasks to Matty and encourage others to do the same. It took a while and repeated explanations, but eventually the ideology started to stick: treat Matty like a junior developer (or, in our case, a member of the Mattermost open source community). Give it a well-scoped ticket with a desired end result. Provide some context up front if there’s a very specific way you want things done. Keep the task size small, and you’ll set Matty up for success.
The challenge was getting someone to use Matty the first time. That first delegation is often make or break, and I spent many hours monitoring the Matty Code agent run from a first time user anxiously waiting for it to succeed at its task.
For those successful first delegations where the result was a PR with a small scoped fix as well as a before and after screenshot of the fix working, that engineer was often converted. In the first weeks of Matty’s rollout, we saw engineers with successful first runs ending up delegating at least three more tasks in the next two weeks. Slowly successes started to reach everywhere, and “throw Matty at it” became part of the vernacular.
Not all first delegations were successful, and these first failures often had a very damaging effect on reinforcing preconceived notions of AI not being ready for this sort of thing. Most engineers that saw a first run failure did not delegate another task to Matty Code for the next four-plus weeks.
When Matty runs are successful, it feels like magic. When it fails, it erodes confidence. A lot of the Matty failures that eroded trust early on weren’t Matty’s fault; they were process issues.
Hard-Won Technical and Process Learnings
The most consistent cause for failure I’ve observed with Matty over the last six months is related to Jira hygiene. It turns out that giving an agent a poorly scoped ticket is just as much a toss-up as giving one to a human.
Matty isn’t an engineer that’s designed to understand requirements through discussion with stakeholders. It’s designed to follow instructions with some agency. Most of the Jira hygiene problems we had that led to poor results from a Matty run fell into a few different buckets.
Tickets without reproduction steps
We all have tickets like this. A staff member working remotely on 3G from their cabin in the remote Canadian woods experienced an issue on the mobile app. They opened a bug, and that bug consisted of a screenshot with the caption, “This is happening and it shouldn’t be.”
To a human engineer with intuition and deep product knowledge, it may be easy to suss out what the reporter means. But an agent can’t. Depending on the chosen model, these types of tickets might get PRs on a spectrum of documentation changes, to massive refactors, to neverending runs attempting to understand what it’s even trying to reproduce.
Tickets without clear expected outcomes
Bugs that are reported with exceptionally accurate reproduction steps often lead to very successful reproduction runs by Matty. But tickets missing an expected result leave too much room for agency where humans should be making the decision.
Bugs delegated to Matty without an expected result in the description were much more likely to end up with the wrong solution (and thus contribute to the anxiety of AI slop).
Obsolete tickets
LLMs are built to generate, and they tend to aim to please. Tell an agent something is broken when it’s not, and some models will end up down rabbit holes burning tokens in attempts to reproduce. With tickets that are over a year old, it’s always good to give a quick repro check ahead of delegation to save the token burn.
Writing better tickets doesn’t just benefit Matty, it benefits humans, too. We’ve spent considerable time in the last month looking through our backlog for poor tickets and rewriting them to avoid the pitfalls discussed above. This allowed us to narrow down hundreds of potential tickets to delegate to Matty.
The Agent’s Environment
The agent’s environment, and its ability to configure it were another major pitfall in the earliest iterations of Matty Code. Areas of the code that require licenses, complex configuration, extensive amounts of backfilled data, or external systems can cause an agent to spin its wheels trying to figure things out.
I frequently (with the help of agents) audit transcripts from Matty Code sessions in our cloud agent environment, looking for cases where an agent was stuck. A good example of this happened within the first few weeks. Matty had been given a short, limited dev license to our Enterprise Advanced product, the highest tier we currently offer at Mattermost, giving access to all features.
While we did communicate that Matty had access to this license, it wasn’t clear from the SKU alone that this license was all powerful. In runs that required a license (often “Enterprise”) for a feature or bug fix, Matty would frequently spend time and tokens looking at the license code, identifying the SKU of the current license, and making sure the feature in question was applicable to the license.
This isn’t a lot of work and it’s not a ton of tokens in isolation, but it gets expensive when added up over hundreds of runs. We fixed this with a single sentence in Matty’s system prompt, and it hasn’t stumbled over this since.
Matty By the Numbers
Matty is actively shipping code at Mattermost, with the oversight of our engineers. Starting at just three tickets in March, Matty is now averaging over 40 tickets per month.
With a conservative estimate of two or three hours per ticket of work for a human (remember that fixed cost I talked about earlier) on average, Matty is rapidly approaching similar productivity to a human hire, allowing the much more expensive, actual human hours to be invested elsewhere.
This number will only grow as we continue to improve the agent, our processes, and our data over time.
The Future
I’ve believed in the vision of Matty Code from the start, and I’m slowly seeing engineers at Mattermost coming around to it as well.
Matty has been a force multiplier. It’s not just taking work off of people’s plates, but building confidence in AI throughout the company, allowing for trust in AI in day-to-day operations that wasn’t there before.
We’ve got a lot planned for Matty in the near future. We rolled out the ability for Matty to work with iOS and Android simulators last week. We’re working to move off of n8n as our orchestrator into a code-driven API, one that Matty itself might be able to maintain someday.
In leveraging agents to work through our Jira hygiene, we have over 200 Matty candidate tickets ready for delegation; we’re aiming to get our backlog to zero and leveraging Matty to keep it there.
As we continue to build out these capabilities, we’ll be able to throw more and more complex problems at Matty and see success on the other end — to the point that Matty can take on more intermediate engineer-level tasks that could span multiple user stories, projects, or PRs.
I’m excited to see where we end up in six months. I’ll be back for another update then!