Adaptive Prompt Embedding Optimization For LLM Jailbreaking
Abstract:Existing white-field jailbreak assaults in opposition to aligned LLMs usually append discrete adversarial suffixes to the user immediate, which visibly alters the prompt and operates in a combinatorial token space. Prior work has prevented immediately optimizing the embeddings of the original prompt tokens, presumably as a result of perturbing them dangers destroying the prompt’s semantic …
Adaptive Prompt Embedding Optimization For LLM Jailbreaking Read More »
