top of page
Search

Mind the Gap - Supporting Pre-Flight, a Benchmark for LLM Assessment in Aviation Operations

  • Writer: Callala Support Team
    Callala Support Team
  • 4 days ago
  • 2 min read

When Alex Brooker at Airside Labs set out to measure how well large language models understand aviation operations, he asked whether Callala would lend a hand. We were glad to. The result, Pre-Flight, is now published, and we are proud to have contributed a little time towards it.



Pre-Flight is an open benchmark of 300 questions drawn from real aviation source material that simply asks whether LLMs can reason correctly about the operational knowledge that airlines and airports depend upon.


Against an informal expert reference of around 95%, the strongest model tested today reaches only 82.7%. Interestingly, over the past 12–18 months, LLM improvement against this test set has been gradual rather than dramatic. What this means is that a persistent gap below expert reliability remains, and that LLMs cannot yet be trusted with the most complex of tasks.


Why we got involved


The first reason is partnership. We have collaborated before, and that earlier work taught me something I have deliberately hung onto ever since: our belief that lasting business partnerships are built on equal business stature and equitable reward. Supporting Pre-Flight was a small way to honour that principle.


The second is curiosity.  For us, specifically about the limits of machine reasoning. We are neither aviation specialists nor pros in agentic AI. Our contribution was contextual, insofar as we understand geospatial logic, much of which fits neatly onto the spatial and temporal constraints and complexities of ground operations. Our interest lies in the cases where an LLM's logic breaks down under real-world constraints.


The third is accessibility. Aviation has a depth of vocabulary, and work like this carries a risk of being used only by insiders. Part of my input was to help frame the benchmark so that people outside aviation can take something useful from it. Basically, anyone aiming to weigh up where AI can, and probably cannot, be trusted.


The wider point


Pre-Flight deliberately keeps to non-safety-critical business operations in aviation, and it is honest about its own limits. Integrity is central to the utility of the benchmark. Looking further ahead, mandating open-ended answers over multiple choice will be central to its development and evaluation over the longer term.


Our thanks to Alex Brooker, Tim Hughes and the practitioners behind the benchmark. It is a genuinely useful contribution, and one we were pleased to support.

 

Useful Links




Graphic by Victoria Beall, adapted with GPT-5.5 Intelligence: High

 

 
 
 

Comments


bottom of page