Language models are often evaluated on curated benchmarks that underrepresent the complexity of enterprise deployments. We introduce HARDEN, a constrained evolutionary search method to adapt the input of existing evaluation cases into more challenging variants while keeping their expected outputs fixed. HARDEN searches...
Aditya Kumaran, Rahul Singhal, Karime Maamari et al.· 0 citations
This work revisits schema linking when using the latest generation of large language models (LLMs) and finds empirically that newer models are adept at utilizing relevant schema elements during generation even in the presence of large numbers of irrelevant ones.
Karime Maamari, Fadhil Abubaker, Daniel Jaroslawicz et al.· arXiv.org· 109 citations· ⚡19
This work systematically pre-training and evaluating on many diverse datasets and analyzes what aspects of the data are most important for building a Tabular Foundation Model (TFM) generalizing across domains to show that the number and quality of tasks one can construct from a dataset is key to downstream performance.
Junwei Ma, Nour Shaheen, Alex Labach et al.· arXiv.org· 4 citations
This paper introduces the first LLM routing approach for Text-to-SQL, which dynamically selects the most cost-effective LLM capable of generating accurate SQL for each query.
Mohammadhossein Malekpour, Nour Shaheen, F. Khomh et al.· arXiv.org· 10 citations· ⚡1
SQLMorph's query mutation and fine-grained metrics support debugging and better align Text-to-SQL evaluation practices with real-world deployments, and introduces a family of execution-level metrics that address the limitations of current binary measures.
Mohammadhossein Malekpour, M. Riahi, Maxime Lamothe et al.· IEEE International Conferenc...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.