Reasoning that holds, however long the chain.

Large language models have two problems that make them unfit for high-stakes reasoning — and neither is fixable with more data. First, they can’t truly reason: their “reasoning” is predicted text, not logical deduction, and they’re strong at spotting patterns but weak at applying rules to new facts — exactly what regulated work demands. Second, they hallucinate, and it’s inherent to how they generate, not a flaw that scaling removes.

01 — Why LLMs fall short

Two problems, one root cause.

It can’t actually reason.

LLM “reasoning” is generated text predicting the next plausible token — not a trace of logical deduction. It holds no internal model of how rules connect, which overrides which, or how a burden of proof shifts. Research bears this out: LLMs are strong at induction (spotting patterns from examples) but systematically weak at deduction — applying a known rule to a novel fact pattern (Cheng et al., UCLA & Amazon, 2024). Regulated reasoning — law, compliance — is overwhelmingly deductive. It is precisely what LLMs are weakest at.

It hallucinates — and scaling makes it worse.

Hallucination isn’t a bug to be trained away. It is mathematically inherent to autoregressive generation (Xu et al., NUS, 2024), and OpenAI’s own analysis traces it to how these models are built and graded. More capable models hallucinate more on factual queries.

Hallucination rate rises with model capability

OpenAI PersonQA benchmark

16%o133%o348%o4-mini

More powerful models do not hallucinate less. Scaling is not the fix.

In a deductive chain, one wrong step collapses the whole conclusion. For regulated reasoning, “usually right” is not good enough.

02 — It’s not the training data

You can’t fix this with a better model.

LLMs predict the next most likely token from statistical patterns. When they “reason,” they mix and match fragments of reasoning seen in training — and for any step not well represented there, the same mechanism produces a confident fabrication: an invented premise, a wrong inference, a citation that does not exist. The failure is structural, not incidental. Independent research groups have reached the same conclusion by different methods — NUS and OpenAI among them: hallucination is an innate property of how these systems generate, not a gap the next larger model will close. In regulated domains it is worse still — more training data sharpens statistical averaging, while the correct deduction often hinges on the edge case (the specific rule for the specific facts), which an averaging system underweights precisely when correctness matters most.

“Hallucination is not a training failure. It is the consequence of asking a statistical system to do deterministic reasoning.”

03 — The stakes

Where “probably” isn’t allowed.

Deductive by nature

A lawyer applies known rules to new facts; a compliance officer applies a regulation to a transaction. This is deduction — the LLM’s weakest mode.

The edge case is the answer

The matter usually turns on the one provision that fits these exact facts. A system biased toward the average answer misses it.

Chains are fragile

One hallucinated step — a wrong rule, an invented authority — and the whole conclusion is unsound. There is no partial credit in court.

04 — Our architecture

Stop asking neural networks to reason.

Reasonex — neuro-symbolic reasoning across legal, banking, insurance, telecom, and healthcare
What it’s good at

Perception · Neural (LLMs)

Reads unstructured input. Extracts facts, handles ambiguity and linguistic variation. Approximate, flexible, probabilistic.

Zero hallucination

Reasoning · Symbolic (deterministic reasoners)

Applies rules to the extracted facts. Traces dependencies. Reaches conclusions that are exact, auditable, and reproducible for the same inputs.

Neither side does the other’s job. The neural layer never reasons; the symbolic engine never guesses. The reasoning that carries consequence is derived by a system that, by design, cannot fabricate it.

Neural

The LLM reads natural language and maps it to structured concepts the engine can act on. It doesn’t reason. It listens.

Symbolic

A deterministic reasoning layer applied to the domain’s governed rules. It surfaces structural patterns, proves logical entailments, and detects conflicts — and because it computes rather than predicts, every result is a mathematically necessary consequence of the rules, not a probable guess.

Deterministic

Gate logic, burden-of-proof allocation, chain traversal, and multi-path resolution over structured domain knowledge. Same input, same output, every time. Fully auditable. Fully reproducible.

How a query flows through Reasonex™User Queryplain languageNeuralperceives languageand intentSymbolic Reasoninglogical deductiondeterministic evaluationValidation Firewallblocks any unverifiedclaim from the outputVerified Outputpathways · firmnessmissing facts · authorities

Perception is neural. Reasoning is symbolic. A Validation Firewall guards the output — so every result is coherent, auditable, and reproducible.

05 — Not fringe

The whole field is moving this way.

Neuro-symbolic AI is where the field is heading. Gartner’s Hype Cycle for Artificial Intelligence, 2025 profiles it as a form of composite AI that, in Gartner’s words, “can augment and automate decision making with less risk of unintended consequences.” In the research literature it’s been called “the third wave” of AI (Garcez & Lamb, Artificial Intelligence Review, 2023) — after symbolic systems and today’s statistical models. The same pattern is already proven in practice: DeepMind’s AlphaGeometry pairs a neural model with a symbolic engine to reach results that hold up to formal checking. Reasonex applies that architecture to regulated reasoning.

Our Chief Scientific Officer, Prof. Ernest Chong, is the pioneer of Algebraic Machine Reasoning (Chong et al., CVPR 2023) — the framework that turns reasoning into exact algebraic computation and exceeded human performance on abstract-reasoning benchmarks.

06 — The platform

Reasonex is the engine. MikeROS™ is the first application.

Reasonex is delivered as the Reasonex Operating System (ROS) — the reasoning layer applications are built on. It is domain-agnostic by design: adding a new regulated domain is a configuration exercise — a vocabulary and a codified rule set — not an engineering rebuild. Domain products inherit its determinism and full audit trail without rebuilding the core. The same engine reasons across:

  • Legal — MikeROS — procedural reasoning over the Rules of Court, live in three jurisdictions: Singapore, Australia (Federal), and New South Wales (NSW)
    Live
  • Claude for Legal — Native integration via Model Context Protocol (MCP)
    Live
  • Network & data centre operations — KaiROS™ — Intent-Based Operations: intent, policy, and change governance reasoned deterministically before anything executes
    In development
  • People-decision compliance — AnthROS™ — every people decision checked against employment law and the organisation’s own HR policy
    In development
Patents
Dual Singapore filings — SG 10202503202W · SG 10202503334Q (patent-pending)
Chief Scientific Officer
Prof. Ernest Chong, originator of Algebraic Machine Reasoning (CVPR 2023)
SUTD ARISE
Selected into SUTD’s ARISE venture-building programme (Venture, Innovation & Entrepreneurship)
Contact

AI that reasons.

hello@reshuffleai.com