Empirical Study of Code Large Language Models for Binary Security Patch Detection
This initial study demonstrates that directly prompting off-the-shelf code LLMs remains ineffective; even advanced prompting strategies cannot compensate for the lack of task-specific knowledge, and fine-tuning proves highly effective, with pseudo-code representation consistently yielding the best performance.