VMTech
Discuss a project

CTF.ae Introduces XRanges for AI Agent Security Evaluation

CTF.ae Introduces XRanges for AI Agent Security Evaluation

CTF.ae has introduced XRanges for AI, an evaluation platform for autonomous security agents that measures what an agent does inside realistic target applications rather than relying on its final report. The platform produces four live signals—coverage, boundaries, exploited vulnerabilities and integrity—and can deploy an isolated multi-container target in about 90 seconds.

The approach has already been exercised at scale. During the 48-hour DEF CON 34 Bug Bounty Village CTF in August 2026, CTF.ae created the Xenoptic target for 545 registered players. More than 850 deployments were monitored using the same four signals now offered for agent evaluation.

Measuring behaviour instead of claims

Autonomous pentesting and bug bounty agents may describe an access-control exploit in a report without establishing whether exploitation succeeded, stopped at an intermediate step or was incorrectly inferred. XRanges for AI is designed to resolve that ambiguity through telemetry collected within each target deployment.

Its benchmark library consists of full applications with business logic, seeded data, background jobs and simulated user activity, built using several languages and frameworks. Each target contains 20 or more injected vulnerabilities, including single-step issues and chains that cross service boundaries. CTF.ae says the targets also include zero-days identified by its researchers and are not present in public training corpora.

Every service emits structured telemetry through OpenTelemetry. CTF.ae says this instrumentation is hand-written for each application by application-security and software engineers, rather than limited to generic HTTP logging. The platform ingests the telemetry for each deployment and updates the scores while the agent is operating.

Four independent evaluation signals

Coverage tracks legitimate business actions, such as registering an account, viewing a job posting or using an assessment feature. These coverage points cannot be reached through an exploit, and the platform lists named actions the agent did not reach.

Boundaries records violations of target-specific rules of engagement, including actions such as deleting hiring content or revoking API keys. A violation is logged when it occurs, with its container and timestamp. Exploited maps a vulnerability to ordered kill-chain phases and records how far the agent progressed using signals from inside the application.

Integrity runs checks every minute to establish whether seed data, service responses and cross-service trust remain functional. A failed check incurs a penalty regardless of cause, helping identify cases where an agent disrupts the environment while pursuing a finding. Although the four measures roll up into a score, the platform presents their breakdown for analysis.

Automation and repeatable comparisons

XRanges for AI can run up to 1,000 deployments simultaneously, CTF.ae says. It records activity in a timeline mapped to business functionality, while engineers can inspect the raw OpenTelemetry stream and query it with regular expressions and attribute filters. If an agent reports an issue outside the vulnerability catalogue, the recorded timeline can help determine whether it was a false positive or an unplanned genuine flaw.

Vulnerabilities can be toggled or patched in a running deployment, allowing retests against the same environment and state. Deployments can also carry metadata such as model name, agent version and prompt variant; grouped results show average and best scores alongside per-vulnerability completion data. The console, API and Model Context Protocol server support batch deployment, collection of coverage and kill-chain progress, and comparisons from CI pipelines or chat assistants.

For businesses developing autonomous security agents, the practical implication is to evaluate repeated runs against verified in-target evidence, including missed functionality, boundary violations and integrity failures, rather than treating an agent’s narrative report as proof of security-testing performance.

#cybersecurity#aiagents#securitytesting#opentelemetry
Open analytics
On the site 0 views
min read 5 23.09.2026
Instagram

CTF.ae Introduces XRanges for AI Agent Security Evaluation

Open the post on Instagram ↗