Insights
Oct 6, 2026

We Keep Getting Surprised. And It’s Fantastic.

A first-hand log of how fast AI changed the way we work at Misfit Labs, and why we never stop experimenting.

Posted by:
Ben Sharpe
Partner

Today was another reminder of how far AI has come for software development.

I'd been running FlyCut, a macOS clipboard manager, whose official release hasn't been updated since 2020. It had this annoying habit of simply disappearing at the wrong time. I lost a password I was updating. Frustrated, I decided to ask Grok why that was happening. It didn't have a good answer. I asked if it would be possible to create a new app that had similar functionality. A one-line prompt. Grok built it, ClipStak, in about five minutes, and migrated all of my saved clips over like it was the most normal thing to do. After a second prompt, it now supported images, a feature I had wanted in FlyCut forever. Mind blown.

Six months ago we would not have believed that an agent could write a complete menu bar application for macOS in Swift. Now it's just another Thursday. 

Misfit Labs started as an expert external dev studio. The team was three engineers: Kyle on iOS mobile, Ant on frontend and Android, and me on backend, plus Scott on design and Joey running product and customers. Over the last twenty months, we've turned into a full-on agentic coding studio, with the expert judgment of every team member still in the loop.

We keep a running list of the AI moments that surprised us. Here's what's on it.

How each of us got in

I wrote my last line of code by hand for a client in January 2025. Cursor and the model of the day became the new normal, though I still reviewed and tweaked every chunk of generated code unless it was boilerplate. Kyle was doing about the same. In February 2025 I built a working app with Cursor and Claude in about thirty minutes.

In early 2025, Scott was openly skeptical of AI, mostly because he wanted less technology in his life overall.

In March 2025, he and I turned his admin Figma screens into a working prototype in about eight hours, using Claude 3.7 in Cursor with a PRD written by Grok. That's when his view started to shift. By February 2026, he was asking which AI tools a designer should use to build apps himself. He has since screenshotted tools like Xcode, Clerk, and Cloudflare and had the AI walk him through the exact setup steps. He gave Claude one icon he'd designed and got back a full matching set of vector icons in minutes, work that used to take weeks. In June 2026, he handed Figma screens to our agent team, and the agents applied the UI changes.

Joey came in through music. In October 2025, he signed up for Suno within minutes of hearing the songs I'd made for my side-project game, The Echo. That was his first step into AI. It led to Lovable prototypes, then Claude Code, and eventually a real mobile app.

Anthony was never a skeptic. He kept pace with Kyle and me the whole way. In March 2025, he dubbed a batch of English videos into Spanish with an AI service for $69, done in minutes. In February 2026, he said one-shot specs made him feel about ten times more productive. A month later, he said the older client app was the last thing he built by hand before AI.

Agents became part of the team

In April 2025, Augment Code nearly finished an Elixir backend for real client work. We went all in on Augment that spring and summer. Then Opus 4.5 shipped in November 2025, and we stopped using Augment entirely. Claude has been the primary model behind our agents since.

In June 2025, an agent reported a feature complete when all it had written were # TODO: Implement this stubs. That's why we now make sure we have behavioral tests for everything, and a model in the Verifier role acts as a strict disciplinarian for the implementors.

Agents had no good way to plug into project management, so in September 2025 I started building CodeBake, an AI-native project management tool rather than one with a chatbot bolted on. By October, agents were picking up CodeBake tasks and posting progress comments as they worked. This work would have taken at least six months to do by hand, and would not have been as feature-rich.

In January 2026 we got on OpenClaw, and Hermes Agent shortly after. In May 2026, Kyle started building our internal agent harness, which now helps us build all of our products.

Old software, rebuilt

I love working on vintage cars. One of my projects has a modern TBI (throttle-body injection) system on it to replace a finicky carburetor. The software used to tune this system was a bear. It was Windows-only, built on .NET 4.5, super-slow, and missing a lot of useful features. I discussed this problem with Claude Opus and we decided to decompile the app and see if we could get it to run ourselves. We got it to run and upgraded it to .NET 8.0 in Visual Studio, greatly improved startup time, added a feature that applies learning-table data to the VE table, and improved the UX. All of that took about a day of prompting.

Then I had Claude write a full product spec from the Windows source and rebuild it as a native Mac app in Rust, including talking to the ECU through the same USB adapter. That took about another week, and the team was amazed. 

This experience is what led me to ask Grok to make a FlyCut replacement today.

Grok will get

In January 2026, I set up a multi-model workflow. Opus proposes three solutions, Gemini and Codex vote, and Opus builds the winner. 

Today we actively use all the frontier models, Claude, ChatGPT, Gemini, and Grok, but for different things. Claude runs our agents. GPT has become a great implementer since GPT-5, and we hand it a lot of build work. I've recently started using Grok Bot to manage some test websites, with pretty good results. We're betting on Grok getting better pretty quickly. I use Gemini when I need to gather information from the web in general or YouTube specifically. And the underdog of research, Manus, is a killer tool.

Why we keep doing this

Being surprised this often is a lot of fun, and it's exciting. That's why we explore and experiment constantly. I've spent over 30 years writing code by hand and loving most of it. I don't write any code by hand now, and I couldn't be happier.

It has also put us in an unusual spot. We've spent twenty months learning how to be seriously productive with every frontier model, not just one, and we've done it on real client work with real consequences. That's a good place to be standing when a company asks how to add agentic development to its software process without blowing up what already works. Helping teams get the most out of that transition is what we're excited about now. If you're curious how to improve your own use of agentic development, CodeBake.ai is the place to start.