What we are
Tool development has reached a pace that outmatches weekly, and perhaps even daily, changes in how digital systems operate. Across a vast spectrum of user experience, this has produced well-meaning debate over social cohesion, media integrity, cyber risk, and more. And, like many other tools, these models can serve both beneficial and harmful purposes, bringing risks and rewards unlike any humankind has faced before.
At EEP, we hope to shed light on these questions. We work toward understanding model deviancy, whether in the form of misalignment with intended tool purpose or maladaptive use against the broader public. We hope that the combination of technical expertise with insights for policy, safety, and security will extend into the very ecosystems where these tools take root.
Full reports are hosted on our Substack; this site carries standalone summaries and supplementaryinteractive visualizations.
What out work entails:
Evidence-based analysis on emerging AI tools and their safety implications;
Technical evaluations of model behavior, fallbacks, and exploitability across deployment contexts;
Benchmarking on the capabilities and risks of available systems;
Technical findings translated into interpretable insights for cybersecurity, social safety, and interested audiences.
AI models exhibit characteristics of a general-purpose technology, yet the risks associated with them are generally difficult to see. Some people use these tools to meet organizational needs, others to automate personal projects. But many possible outcomes remain unexplored, compromising our ability to control systems when deemed necessary. EEP intends to derive transparency from the “black box” systems used both systemically and individually, reporting on their inner complexities in a holistic manner.
How we accomplish what we set out to do:
- Test deployment parameters and interaction behaviors over thousands of repetitive rounds using a proprietary, in-house software (i.e. machine-learning security operations);
- Use regression testing to monitor changes across model updates and replacements;
- Link payloads, regex patterns, text logs, and other technical artifacts to our findings for transparency;
- Use visualizations to enhance accessibility and shareability of reporting;
- Study both technical and psychological vectors of exploitation, including those mediated by natural language;
- Connect findings to current affairs and industry standards through interviews with industry professionals.
How findings are classified
Each report is tagged against two established taxonomies, so a reader can connect a specific result to the broader threat landscape and to their own controls.
OWASP LLM Top-10
The community-standard catalogue of the most critical vulnerabilities in LLM applications.
MITRE ATLAS
A knowledge base of adversarial tactics and techniques against AI-enabled systems.
Who we are
A small group doing careful, repeatable work.