HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference
HE-Guardrail is proposed, a framework that evaluates guardrail mechanisms entirely over encrypted data and homomorphically controls whether the target-model response is returned to the client, with distinct security-efficiency-utility trade-offs.
Byeongseo Min, Y. Lee, Young-Sik Kim et al.
· 0 citations