
Decision Models versus Reasoning Models: An Exploratory Evaluation of Perplexity's Decisions API
Introduction
A recent video by IBM Technology inspired me to find out more about the trending Decisions API. When I dug a bit deeper, I found that several alternatives had been released back to back by Perplexity, OpenAI and TypeSafe AI, and that they work somewhat differently from conventional models. Instead of returning chunks of text as intermediate reasoning and final outputs, they return a decision with a probability.
(Source: Perplexity Decisions API documentation)
A problem I'd already run into
In one of my positions, I worked on processing policy documents in Scotland, where I found a big limitation in conventional LLMs. They sometimes misinterpret outdated policies, policies from other documents and irrelevant ones as ones proposed by the processed document, which is a critical error. Part of this was poor architecture design with insufficient inputs. A long policy document, together with the lengthy criteria we had to provide as the system prompt, was often too much for the context window of the models we used.
This is now largely solved in the latest models, which can take the full document and criteria at once. However, as my experiment shows, even with the full document as the input, the problem still exists to some extent. This is certainly not hallucination; it is a weakness in how these models separate what a document proposes from what it merely mentions.
Trying the Decisions API as a guardrail
When I heard about the Decisions API, I wanted to give it a try. Given how it operates, I thought it could at least fact-check findings very cost-effectively, so that we could use it as a guardrail for the outputs of the main reasoning AI models. I ran this as an experiment and obtained interesting results, especially on cost, consistency and its ability to flag its own uncertain answers.
You can have a look at it in the attached paper, previewed below. Please note that it was prepared with the use of AI and has not been peer reviewed for accuracy.
(Preview of the paper's first page, including the abstract and results summary)