Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face

Training Frontier Models to Out-Think Hackers — Uri Rolls, Arithmetic & Thom Wolf, Hugging Face

Source: YouTube · AI Engineer · published Jul 24, 2026 · 17:27

Cybersecurity
No ratings yet Log in to rate
Transcript Available
Description

The speakers introduce "Masov," a rigorous benchmark evaluating AI models' ability to reason through complex, chained logic vulnerabilities in access control, arguing that open-source models are essential for defending against rapidly evolving, multi-target cyber attacks 2:15.

Key Takeaways:
• The benchmark tests models in a black-box environment where they must exploit logic-based vulnerabilities across chained microservices without code access, requiring dynamic world modeling to navigate state changes 7:30.
• Current state-of-the-art models struggle with logical leaps; only GPT-5.5 has successfully solved the benchmark, highlighting a critical gap in reasoning capabilities for complex exploitation chains 12:25.
• Traditional defensive stacks are failing because attackers can now target multiple systems simultaneously, necessitating faster, model-driven defense mechanisms that operate at scale without heavy human intervention 5:20.
• Success requires high-quality, human-curated data to teach models dynamic reasoning rather than simple pattern matching, derived from real zero-day vulnerabilities in open-source software 8:30.
• The future of cybersecurity relies on a collaborative ecosystem of strong open-source models to ensure defensive speed and capability outpace attackers through shared post-training data and collaboration 14:50.

The presentation underscores a critical shift in cybersecurity economics, urging the community to leverage open-source AI to rebuild defensive stacks capable of operating at the speed of modern, automated attacks.

Sources:

  • 2:15 Introduction to the bench

Generate CPE Credits

Generate a professional CPE document from this video's transcript.

Estimated credit: 0.5 CPE hours

Estimate uses the video runtime (1 hour ≈ 1 CPE, rounded to the nearest 0.5, minimum 0.5, maximum 2.0). The final amount can be lower after review, never higher.

Topic: Cybersecurity. Commonly maps to: Security and Risk Management, Security Operations. Exact CISSP domains are assigned during generation.

CISSP Domain Mapping
Learning Objectives
Self-Assessment Questions
PDF Export Ready

Free account. One generation at a time, with a daily limit.

CPEBuddy is independent and not affiliated with or endorsed by ISC2, ISACA, or any certification body. Exports are formatted for common CPE submissions; acceptance is at your certification body's discretion.

Watch on YouTube

Transcript Preview

First 800 characters of the transcript

Hello everyone. Okay, thanks for showing up at this data quality. So as you saw, we probably talk a little bit about other things than data quality. and uh and first I actually tell you about why I'm very excited about this talk and why actually I accepted to uh to come present it with Fury. Um there's two reason for that and the first reason is I think and what you will see today is that cyber security is a much wider field for AI a much wider playing field and exploration field than you might think. And in particular what we'll show you today is uh a benchmark that arithmetic and Yuri has been developing and I've been doing a little bit advising which I think is even close to things like arc AGI 3 for people who has been following progress around AGI in which that um by which I mean that…