Buying a text-to-SQL platform is an exercise in seeing past the demo, because every vendor’s demo works. The model generates clean SQL against a tidy schema, the answer appears in seconds, and the room is impressed. The differences that decide whether the platform survives contact with your warehouse, your permissions, and your auditors are invisible in that moment, and they have to be extracted deliberately. The twelve questions below are designed to do the extracting. Each comes with the answer a serious platform should give and the evasion that should end the conversation.
Accuracy and grounding
1. Will you run the evaluation on our warehouse, at its full size, before we sign?
The only accuracy that matters is accuracy on your schema, with your ambiguity and your table count. A confident vendor agrees to a structured trial against your environment using your questions. A vendor who insists the demo environment is representative is telling you their accuracy does not survive scale, whether they know it or not.
2. Where does the system get its business definitions, and who certifies them?
The right answer describes a semantic layer: a governed store of metric definitions, blessed tables, and join paths that the business signs off and the model must use. The wrong answer is any version of “the model understands your schema,” because schemas do not contain business meaning, and a system inferring definitions per query will contradict itself across users and phrasings.
3. What happens when the system does not know what a term means?
Enterprise-grade behaviour is to ask for clarification or decline, and to record having done so. If the demo never shows the system declining anything, ask to see it. A platform that always produces an answer is a platform that guesses, and guessing is where silently wrong numbers come from.
4. How does accuracy improve after deployment, and does improvement persist?
Look for a concrete learning mechanism: analyst corrections that flow into the semantic layer and hold for every subsequent user, which is how QuaerisAI’s Smart Semantic Layer compounds. Be wary of answers that reduce to “the models keep getting better,” because model upgrades do not teach the system your fiscal calendar, and per-user chat memory that evaporates between sessions is not organisational learning.
Governance and security
5. Whose permissions does a generated query run under?
The only acceptable answer is the asker’s, enforced at query time, inherited from your identity provider. Many products connect through a broad service account and filter results in the application layer, an arrangement that leaks the first time the model generates a query nobody anticipated. Ask the vendor to demonstrate two users with different entitlements asking the same question and receiving appropriately different results.
6. Show us the complete audit record for one query, live.
Phrased as a demand rather than a question, deliberately. The record should link, as one retrievable chain, the asker and their role, the prompt as typed, the definitions it resolved against, the generated SQL, the tables touched, the model used, and the result. If producing this requires stitching separate logs together or an engineering ticket, the platform has logging, not lineage, and your auditors will notice the difference.
7. Where does our data travel, and what reaches the model vendor?
You are listening for a warehouse-native architecture in which queries run where the data lives, with no bulk data movement into the vendor’s environment, and a precise account of what appears in model prompts. Vague reassurance on this question is disqualifying in a regulated industry, because your CISO will ask it with subpoena-grade precision later.
8. Which language models can power the platform, and can we change our choice?
Model choice is a compliance and procurement question, not a technical preference. A BYOM architecture lets you select among providers such as Anthropic, OpenAI, Google, and Meta, and switch as your requirements, your regulators, or the model market change. A platform bound to a single model vendor transfers that vendor’s roadmap risk and pricing power directly onto you.
Scope and operations
9. Can it answer questions that span structured and unstructured sources?
A growing share of real business questions joins warehouse data with what lives in documents: contracts, policies, claims files, reports. A platform limited to tables answers half the question and leaves the other half to shadow tools. Ask for a live demonstration of a single query drawing on both, with the same governance applying to each side.
10. What does deployment actually require, in time and in our people’s hours?
Get specific: weeks or quarters, and how much semantic modeling must exist before first value. Architectures requiring exhaustive upfront modeling of the warehouse impose a tax measured in months; platforms that start from a certified core and learn outward reach usefulness far sooner. QuaerisAI deployments, based on QuaerisAI published materials, are measured in days rather than quarters, and any vendor should be held to stating their number equally plainly.
11. What breaks when our warehouse changes?
Warehouses evolve continuously: new tables, renamed columns, migrated sources. Ask how the platform detects and absorbs schema change, and what the failure looks like when a certified definition points at something that moved. The poor answer is silent breakage discovered by users; the good answer involves detection, flagged definitions, and a governed correction path.
12. Which of your reference customers looks like us, and can we speak to them?
Like you means your industry’s regulatory posture, your warehouse scale, and your user population, because text-to-SQL success is environment-dependent in exactly those dimensions. A vendor without a comparable reference is asking you to be the reference, which is a discount conversation, not a list-price one.
Using the checklist
The twelve questions sort into a simple scoring posture. Questions 2, 5, 6, and 8, the semantic layer, query-time permissions, the audit chain, and model choice, are structural: a platform missing any of them cannot be remediated by configuration, and in a regulated industry each is individually disqualifying. The remainder calibrate cost, speed, and fit. Run the structural four in the first vendor meeting, before anyone invests in a trial, and reserve the trial itself for question 1, executed on your warehouse with a gold-standard question set your analysts have answered independently. An evaluation run in this order costs days and routinely saves the year that a failed deployment consumes.
Frequently asked questions
Should this checklist apply to tools bundled with our warehouse or BI suite?
Yes, with particular attention to questions 5, 6, and 8. Bundled tools inherit an assumption of safety from the platform they ship with, but bundling does not create query-time permission enforcement or prompt-level lineage where the architecture lacks it, and bundled tools are usually bound to their vendor’s model choices by design. The checklist is vendor-neutral precisely because the gaps are not.
How long should a proper evaluation take?
With the structural questions asked up front, a decisive trial fits in one to two weeks: connect to the warehouse, load the certified core definitions, run a fifty-question gold-standard set, and review the audit records. Evaluations that stretch to months usually indicate either an architecture requiring heavy upfront modeling or an evaluation with no defined pass criteria, and both are findings in themselves.
Who should be in the room for the vendor sessions?
The data team owns questions 1 through 4 and 9 through 11; security and compliance own 5 through 8; and the business owner of the pilot domain should witness question 3, because how a system behaves at the edge of its knowledge is the clearest single preview of whether their team will trust it.

