← back to blogs

I completed Flare-On in 91 minutes

Originally published on Medium

I wanted to see how good frontier models were at reverse engineering. I did not expect this.

Illustration of the agent setup

Illustration of the agent setup.

A Reddit post sent me down this rabbit hole. Someone in cybersecurity described a colleague who usually needed days to finish Flare-On. With Claude, he apparently finished in two hours.

The author has since deleted the post. I saved a screenshot, but I wanted to test the claim myself anyway.

The Reddit post that prompted the experiment. The author has since deleted it.

The Reddit post that prompted the experiment. The author has since deleted it.

My first attempt was with DeepSeek V4.1 Flash. It got through the introductory task, then spent the rest of the night stuck on the next one. The two-hour story seemed a lot less believable by morning.

The next day, I tried a different setup. Sonnet 5.5 as the orchestrator, with GPT-6 Astra, GPT-6.1 Sol, and Opus 5.5 as workers. I gave it a cURL and told it to find open-source tools, coordinate the workers, and solve the CTF.

The agents solved the remaining eight challenges and submitted the flags themselves, all in 91 minutes. My account, neutr0n, finished all nine and sits at 401st place.

During the run, I didn’t even open the browser. I wasn’t choosing tools, explaining the binaries, or uploading answers. This was a casual experiment, and the agents went ahead and completed the whole fucking thing.

How unusual is that?

Flare-On is an annual reverse-engineering CTF run by Google’s Mandiant FLARE team. You receive programs and files, figure out their hidden behavior, and recover answers called flags. You generally don’t have the source code. The programs are deliberately made difficult to understand.

Finishing has been a serious achievement. In 2025, 313 people completed it, out of 4,139 registered users. About 7.6 percent.

The 2025 honor roll puts the first finisher at one day, three hours, and 47 minutes after opening. Tenth place took more than three days.

I checked this year’s top ten against their public solve timestamps. A hypothetical finish 91 minutes after opening would have slotted into sixth place, assuming the introductory task was also accounted for. Tenth place finished in about 128 minutes.

I actually finished on September 30, several days after opening. That is why my rank is 401st. The pace itself was already competitive with the top ten.

My public profile. The rank reflects when I finished relative to the contest opening.

My public profile. The rank reflects when I finished relative to the contest opening.

They handled the whole job

The orchestrator researched approaches and split up the work. The workers saved their findings in shared files, wrote scripts, and checked each other’s results. Sonnet followed their progress and submitted the answers.

They made mistakes too. Workers duplicated some work. A brute-force job overloaded the machine. Some assumptions turned out to be wrong, and another worker’s independent check caught them. They recovered without me personally solving the problems.

Workers compared findings and tested each other’s assumptions.

Workers compared findings and tested each other’s assumptions.

Only one refusal

Opus 5.5 refused one task. None of the other models refused any of the tasks in this run. The work was reassigned, and they carried on.

I’ve encountered cyber guardrails in other CTF experiments, especially web-security work. Here, reverse engineering and cryptography barely interrupted the agents at all.

A legitimate CTF solve is not proof that safety controls failed. Authorized security research is allowed. But a single provider’s refusal did very little to restrict what this group of models could accomplish.

One model refused. Other models continued.

One model refused. Other models continued.

I’ve participated in CTFs before, and I build agent systems. I know enough to appreciate what these agents did. I also know I couldn’t have done the same work myself that afternoon.

Yet I could ask for it, and get it.

I started because a Reddit post sounded hard to believe. Now I’m wondering how many other people can do this, and what they’ll ask for next.

Flare-On 13 is still running and ends on October 23, 2026. I’ve left out flags, solutions, and challenge details. After it ends, I’m happy to share the full technical writeup and agent setup.