A benchmark can show 99% refusal while your agent is 76 to 89% breakable. The gap is structural: static tests only check the layer that is already defended. An attacker who adapts across turns walks past it.Um benchmark pode mostrar 99% de recusa enquanto seu agente é 76 a 89% quebrável. A lacuna é estrutural: teste estático só checa a camada já defendida. Um atacante que se adapta turno a turno passa direto.
Our preprint formalizing the two-layer model, the seven attack classes, and the evaluation protocol this audit runs. Introspective Vulnerabilities in LLM Alignment · Manuel G. Galmanus · Bluewave AI Research.Nosso preprint que formaliza o modelo de duas camadas, as sete classes de ataque e o protocolo de avaliação que esta auditoria roda. Introspective Vulnerabilities in LLM Alignment · Manuel G. Galmanus · Bluewave AI Research.