Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Paper • 2608.08722 • Published 14 days ago • 7
A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization Paper • 2608.08156 • Published 15 days ago • 10
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure Paper • 2608.08722 • Published 14 days ago • 7
A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization Paper • 2608.08156 • Published 15 days ago • 10
Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas Paper • 2605.30003 • Published May 28 • 2
Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas Paper • 2605.30003 • Published May 28 • 2
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon Paper • 2605.09708 • Published May 10 • 5
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon Paper • 2605.09708 • Published May 10 • 5
Discovering Agentic Safety Specifications from 1-Bit Danger Signals Paper • 2604.23210 • Published Apr 25 • 4
Discovering Agentic Safety Specifications from 1-Bit Danger Signals Paper • 2604.23210 • Published Apr 25 • 4
STEM Agent: A Self-Adapting, Tool-Enabled, Extensible Architecture for Multi-Protocol AI Agent Systems Paper • 2603.22359 • Published Mar 22 • 4
Cooperation and Exploitation in LLM Policy Synthesis for Sequential Social Dilemmas Paper • 2603.19453 • Published Mar 19 • 6
Cooperation and Exploitation in LLM Policy Synthesis for Sequential Social Dilemmas Paper • 2603.19453 • Published Mar 19 • 6
Multimodal Models 🔀 Collection A collection of multimodal models developed by the Komorebi AI team • 3 items • Updated Sep 23, 2025 • 2
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement Paper • 2507.18742 • Published Jul 24, 2025 • 6