Statutory-Domain Asymmetry in LLM Compliance Guardrails: A Cross-Model Study of Hard Refusal Failure and Soft Gating Collapse in Enterprise HR Contexts
Abstract
We evaluate the resilience of two frontier LLMs—Gemini 3.1 Flash and Claude Sonnet 5—against executive-authority pressure to draft unlawful HR documents across ten U.S. employment-law domains: ADEA, NLRA, ADA, Title VII, FLSA, FMLA, OSHA §11(c), ERISA §510, SOX §806, and PWFA/PDA. Each scenario-model pairing in the original experiment represented a single test run (N=1 per cell), with limited informal replication on selected cells. The original findings are therefore preliminary and illustrative rather than statistically powered benchmark estimates. The experiment used statute-specific system prompts and a three-turn escalation structure: Executive Directive, Pretextual Framing, and Operational Weaponization. The original experiment identified two broad failure classes. Hard Refusal Failure occurred when a model generated an operationally usable prohibited instrument despite recognizing the underlying legal or ethical problem. Soft Gating Collapse occurred when initially stated safeguards weakened under ambiguity or reframing. Operational failures were further characterized through three mechanisms: Proxy Laundering, Disclaimer-Shielded Capitulation, and Chain-of-Thought Safety Decoupling. In the original N=1 experiment, Gemini 3.1 Flash exhibited Hard Refusal Failure in seven of ten domains, while Claude Sonnet 5 resisted Hard Refusal Failure in all ten directly tested domains and two additional attack vectors. In a separate test using a non-canonical system prompt, Claude exhibited Soft Gating Collapse under Ambiguity Deflection. A follow-up FLSA experiment initially suggested that demographic framing explained Gemini’s behavior. Replication did not confirm that interpretation and instead indicated model version as the stronger explanation: Gemini 3.1 Flash collapsed regardless of framing across eight runs, whereas Gemini 3.6 Flash held regardless of framing across three runs. A subsequent API follow-up evaluated 300 independent conversations across the same ten statutory domains, two scenarios, five repetitions, and three Gemini model/thinking configurations. Sixteen operational failures were identified, all within ADEA; the other nine domains held across all 270 conversations. Within ADEA, observed failure rates were 60% for Gemini 3 Flash at LOW thinking, 50% for Gemini 3 Flash at MEDIUM thinking, and 50% for Gemini 3.1 Flash-Lite at HIGH thinking. Because model identity and thinking configuration were not fully crossed, these conditions do not isolate a causal thinking-level effect. The original seven-of-ten-domain result did not replicate in the larger follow-up. The supported conclusion is therefore narrower: refusal language alone is not sufficient evidence of safe behavior, and an ADEA-specific proxy-laundering vulnerability persisted across the tested Gemini configurations. Model-version differences, prompt-provenance limitations, small within-cell samples, and the exploratory character of the original experiment constrain broader generalization.