Skip to content

Honeypot experiment

A honeypot run by a language model

Two decoy servers on the public internet, built and run for a whole month by a language model. What was thrown at them, how it was measured – and all the data, including a complete record of the agents’ work.

Data updated: 2026-10-03 10:31:16 UTC+02:00

Report a problem

The run in numbers #


Time window
2026-08-21 – 2026-09-16
Source data
honeypot-ports-aggregated-20260922.csv, timeline-aggregated not available yet, samples-aggregated not available yet
Script
to be published

What this is #


From 2026-08-21 to 2026-09-17 (28 days), two virtual servers with public IP addresses ran at the same provider. Each hosted a honeypot – a decoy system that poses as a vulnerable server and records everything anyone tries to do with it. Both got the same brief, yet they ended up very different: srv3 offered about 46 specific services, while srv4 answered on all 65,534 TCP ports. The goal was to find out which ports attackers care about most.

The honeypots were not built by a person. Both were deployed and operated for the whole month by a language model – Opus 5 on one server, Fable 5 on the other. The models got root, a brief and nobody to advise them; even how many ports to open was their call. Checks happened about twice a week, each in a fresh chat with no memory of the previous ones. The complete record of what they did, mistakes and outages included, is being published.

Results will be added once the run has ended and the data is processed. Until then the numbers in this section are placeholders; the methodology and the description of limitations already apply.

Recording #


Where to next #