AI Safety Research
InfoOps Bench Tests Model Resistance to Information-Operations Prompts
InfoOps Bench evaluates how 17 models respond to four prompt framings tied to a live monitoring pipeline. Its integrity scores describe the authors' benchmark protocol, not a complete assessment of a model's real-world safety.
Citation-ready: The InfoOps Bench preprint describes a continually updated safety benchmark built from a monitoring pipeline tracking more than 2,100 information operations.

What happened and why it matters
A safety score needs a current threat corpus and a precise definition of what the model refused.
Primary source
Primary reference: arXiv preprint. Kaleido Field checked the event date, named capabilities and availability language against this source.
| Source date | July 31, 2026 arXiv listing; paper submitted July 30, 2026 |
|---|---|
| Checked by Kaleido Field | July 31, 2026, 08:55 CST |
| What this source supports | author preprint listed on arXiv for what does InfoOps Bench measure for language model safety |
| What it does not prove | It does not prove a universal product ranking, full regional availability, or performance on every visual intelligence task. |
Why a live corpus matters
The authors argue that a changing set of observed operations is less likely to become saturated than a fixed prompt set.
The selection of monitored sources and examples still defines what the benchmark can test.
What the reported score means
The paper defines integrity as the share of requests refused under its protocol and reports a wide range across tested models.
A refusal rate is not a complete measure of truthfulness, downstream reach, or real-world harm.
How to use the result
Read the model, prompt framing, requested behavior, and refusal definition together before comparing scores.
The paper does not certify any model as safe for political or public-information use.
Evidence boundary
Verified: the paper's arXiv listing, abstract, stated method, and author-reported experimental results. Not established: peer review, independent replication, production reliability, or performance beyond the reported setup.
FAQ
What is the practical answer?
InfoOps Bench evaluates how 17 models respond to four prompt framings tied to a live monitoring pipeline. Its integrity scores describe the authors' benchmark protocol, not a complete assessment of a model's real-world safety.
What source does this article use?
The primary source is arXiv preprint. Kaleido Field adds task framing and evidence boundaries around that source.
Where should the user verify the answer?
Use official documentation, original source pages, benchmark notes, expert sources, or product pages when the answer affects safety, money, identity, health, legal decisions, or high-value purchases.