An independent publication evaluating how frontier tools and systems behave under real and simulated conditions, combining reproducible technical testing with safety, security, and policy recommendations.
Global, sociotechnical research of emerging technology.
We study the safety and security vulnerabilities of technological tools and systems, and publish the methods alongside the findings.
Driven by the desire to explore and document the capabilities of contemporary machines (including but not limited to AI) the project aims to improve the design, deployment, and utility of digital infrastructure by informing cybersecurity and social safety research. Our means are tangible, technical parameter assessments of programs available on the market; our end vision is greater transparency for all stakeholders.
Each links to the full report on Substack.
How AI code assistants confidently hallucinate non-existent packages, and how attackers exploit these package hallucinations to pre-register malicious payloads.
An empirical study of indirect prompt injection via RAG poisoning in a SOC analyst scenario, finding that retrieval is the load-bearing security control.
Reporting on industry knowledge alongside international governance frameworks.
A gallery of experiments, educational resources, and interactive environments where we make technical concepts accessible.
Watch a RAG knowledge base light up which documents a query pulls, identifying the poisoned one.
Open →Companion toBlind SentinelAn AI-generated import block resolves to a malicious package an attacker registered against a hallucinated name.
Open →Companion toSubtle Ways to Tell a LieAgents rendered as ants across three environments.
Open →Generative simulationFree of charge, published on Substack.