Skip to content
FIELD NOTE

Two QANTIS IBM Heron papers, one increasingly precise research question.

The first QANTIS preprint mapped a broad quantum decision platform across POMDP belief conditioning and multi-target data association. The second deliberately narrowed the claim to a calibrated rare-evidence service that returns an ordinary posterior to a classical planner.

July 20, 202615 min readNeura Parse Research
QANTISIBM Heronquantum researchPOMDPamplitude amplificationMTDA
QANTIS belief-update service diagram with sequential prior and observation inputs, a central calibrated quantum evidence stage, and posterior outputs for a classical decision pathProduct interface · illustrative values

First-paper reported campaign count

Reported usable-shot yield

Second-paper high-shot max Hellinger

Public arXiv v1 preprints

Abstract

QANTIS moved from a cross-backend campaign reported as 45 experiments to a controlled sequential belief-update service. Reading both papers together shows what survived hardware testing, what was narrowed, and what still needs stronger evidence.

Gap map

The two papers move from platform breadth to a planner-facing inference contract while preserving classical control around the quantum step.

01

Paper 1: map the platform

  • 45 reported campaign experiments on three IBM Heron QPUs, not independent replications
  • POMDP belief conditioning and closed-loop Tiger runs
  • FPC-QAOA experiments for MTDA QUBO instances
02

What the first study exposed

  • Rare-event amplification can increase usable-shot yield
  • Small sequential belief loops can remain close to exact Bayes
  • MTDA optimization reaches a current-device feasibility boundary quickly
03

Paper 2: isolate the service

  • Prior and observation model enter from classical software
  • Heron estimates rare-event evidence under calibration
  • An ordinary posterior returns to a classical planner
04

Claims still outside scope

  • No total-runtime or wall-clock advantage result
  • No replacement of policy selection or action execution
  • No peer-reviewed or production-scale validation yet
01Research arc

The foundational QANTIS preprint asked a broad question: can a hybrid quantum-classical platform support decision problems under uncertainty on present IBM Heron hardware? It covered POMDP belief conditioning, a closed-loop Tiger example, and multi-target data association formulated as a penalty-constrained QUBO and tested with fixed-parameter-count QAOA. That breadth was useful because it placed inference and optimization in one architecture, but it also joined circuit families with very different resource requirements.

The July preprint asks a more disciplined systems question. It removes policy selection and action execution from the quantum boundary, gives a classical planner responsibility for both, and evaluates the quantum processor as a calibrated rare-evidence estimator inside one repeated Bayes update. The input is a prior and observation model; the output is an ordinary probability distribution. That narrowing makes failure, fallback, and validation easier to specify.

Taken together, the papers describe research maturation rather than a straight performance race. The first establishes a landscape and reveals where current hardware bends. The second chooses one shallow, decision-relevant primitive from that landscape and tests whether it can be reused sequentially without corrupting the planner-facing posterior.

broad decision platform -> measured hardware boundaries -> calibrated evidence service -> ordinary posterior -> classical decision loop
02First paper

arXiv:2603.00785 reports a campaign count of 45 experiments across three IBM Heron quantum processors. That aggregate spans different workloads and conditions and is not 45 independent replications of one result. Its rare-observation POMDP experiment increased the accepted-event probability from 0.179 to 0.907. In the reported shot budget, that changed usable accepted samples from 1,463 to 7,429, a 5.1x usable-shot yield, while the posterior reached Hellinger distance 0.0015 from the exact reference. The compiled circuit was shallow at ISA depth 18.

That 5.1x result is a sampling-yield observation, not a wall-clock speedup. The paper does not demonstrate that submitting, queueing, executing, mitigating, and decoding the quantum job is faster than a classical update. Its value is narrower: when accepted evidence is sparse, amplification can redirect a fixed measurement budget toward informative outcomes while preserving the resulting posterior in this case.

The closed-loop Tiger runs add a systems check. The reported T=8 trajectory reaches a maximum Hellinger distance of 0.0149, while the T=4 replications range from 0.0095 to 0.0169. They support the idea that a hardware-derived posterior can be fed back as the next prior over a short horizon, but they do not yet establish stability for arbitrary observation sequences, state spaces, or reward models.

Read 5.1x as usable accepted-shot yield in the reported experiment, never as a 5.1x end-to-end speedup.
QANTIS claim-boundary map placing a quantum evidence core inside calibration, hardware behavior, classical baselines, results, provenance, and review layersTwo-paper evidence map
FIG · RESEARCH EVOLUTION — The second QANTIS paper does not erase the first; it turns one part of the earlier platform into a narrower, more testable service contract.
03Optimization boundary

The first paper's multi-target data association track translated association constraints into QUBO instances and used FPC-QAOA to favor feasible assignments. At 11 variables, the reported solution quality was about 64.1% with a 3.3% spread and an execution time around 180 seconds. In the same study, the Hungarian and GNN baselines reached 100% reported solution quality in less than 0.1 milliseconds for the tested comparison.

At 19 variables, reported quantum solution quality fell to about 20.4%. That decline matters more than an isolated good sample because a tracker needs reliable, timely assignments. On these small tested instances, the classical methods were decisively more practical; QANTIS did not replace them. The paper infers an empirical NISQ boundary around 11–15 variables for this formulation and implementation; the directly reported hardware points that bracket the decline are 11 and 19 variables.

This is a useful research outcome. It separates a quantum formulation that is scientifically testable from a system that is operationally competitive. The follow-on paper responds by leaving MTDA optimization outside its primary validation and concentrating on a shallower inference primitive whose output can remain compatible with classical planning software.

  • 11-variable FPC-QAOA: about 64.1 ± 3.3% reported solution quality, roughly 180 seconds in the reported setup.
  • Tested Hungarian and GNN baselines: 100% reported solution quality in less than 0.1 milliseconds.
  • 19-variable FPC-QAOA: about 20.4% reported solution quality, showing rapid degradation.
  • Supported conclusion: a measured current-hardware boundary, not quantum optimization advantage.
04Second paper

arXiv:2607.06760 replaces the broad platform benchmark with a controlled sequential Tiger case study. Fixed-point amplitude amplification is applied at every listen step so the service does not depend on a Grover guard that must skip concentrated-belief regimes. The principal output is not the amplified amplitude by itself; it is the normalized posterior returned after each hardware update.

The reported 8-step high-shot run uses 32,768 shots per step and reaches a maximum Hellinger distance of 0.009. The 12-step primary run reaches 0.021, while 20-step and 32-step trajectories serve as supporting controls. These results indicate that shallow repeated belief updates can stay within the reported accuracy band on the tested Heron configurations.

Resource context changes the reading. A 10,000-shot matched FPAA run reports a maximum Hellinger distance near 0.033. The guarded Grover comparison used an adaptive 8,000-to-16,000-shot allocation averaging around 10,000, so the table does not prove strictly equal-budget FPAA superiority. A constant-shot comparison remains a clean follow-up experiment.

The 0.009 high-shot result and the approximately 0.033 matched-shot result answer different questions and must remain visibly separated.
05Calibration envelope

The second paper adds controls that explain where the sequential service can and cannot be trusted. A boundary-aware BIQAE protocol addresses estimator fragility near amplitudes zero and one. In reported paired Pittsburgh checks, absolute error at target amplitude 0.01 moves from 0.6317 for the baseline to 0.00224 after calibration; at amplitude 0.95, it moves from 0.4890 to 0.00773. These are same-context calibration results, not a universal correction curve.

A separate rare-event circuit family tests evidence as small as one in a million. The reported accepted probability is 0.9718 against an analytic value of 0.9714, described as 971,832x logical amplification. Because the instance compiles to at most six single-qubit basis gates, the measurement matches the predicted amplified acceptance in a deliberately shallow one-qubit case. It is not a 971,832x runtime claim and it does not validate deep multi-qubit inference.

Scaling probes reinforce that distinction. An optimized four-state sweep reports six of six cases below Hellinger distance 0.05, while a deeper synthesis stress path reaches the strict band in only three of six cases, though five of six preserve the MAP state. State count alone is therefore a poor readiness measure; compiled depth, routing, estimator visibility, and mitigation behavior govern the operating envelope.

06What carries forward

Across both papers, the useful architecture is hybrid. Quantum routines are invoked for a bounded inference or optimization primitive, exact or strong classical methods define the reference, and conventional software owns planning, policy, and action. This makes the quantum result replaceable: if calibration, depth, or budget gates fail, the system can return to a classical estimator without redesigning the decision loop.

The newer paper improves claim hygiene by dividing results into primary sequential evidence, supporting controls, exploratory scaling probes, and out-of-scope system claims. That hierarchy should also govern product communication. A hardware result can be meaningful without being a speedup, and a posterior can be decision-consistent in one reward rule without proving autonomous-system safety.

The public GitHub repository provides the Community Edition software surface, while the papers supply the formal experiment narratives. Independent review would be stronger with one versioned campaign bundle connecting circuit sources, transpiler settings, calibration snapshots, job identifiers, raw counts, derived tables, and plotting code for every reported row.

07Next tests

The strongest follow-up is not a larger amplification headline. It is a preregistered comparison that fixes shot budgets, records queue-inclusive and queue-exclusive timing, repeats across calibration windows, and tests a classical estimator with equivalent statistical confidence. The sequential schedule should include adversarial observation paths that drive beliefs repeatedly toward both probability boundaries.

A second line should connect posterior error to more than one immediate Tiger threshold. Decision regret under multiple reward models, posterior calibration, failure detection, and fallback frequency would show whether distribution fidelity translates into robust planner behavior. Larger-state experiments should be admitted only when their compiled depth and measurement budget remain inside declared gates.

Both QANTIS records are public arXiv v1 preprints as of July 20, 2026. They are valuable, inspectable research records, but they have not yet crossed peer review. The honest synthesis is therefore balanced: hardware evidence has become more precise and more useful, while runtime advantage, broad scale, independent reproduction, and production autonomy remain open.

Practical takeaways

01

Read the first paper as a broad hardware map and the second as a narrower sequential inference case study.

02

Treat 5.1x as usable-shot yield, not speedup, and 971,832x as logical amplification, not runtime advantage.

03

The MTDA experiments identify a current-device boundary; the tested classical baselines remain far more practical.

04

Keep the 32,768-shot FPAA result separate from the 10,000-shot matched control.

05

The stable interface is an ordinary posterior returned to classical planning, policy selection, and action execution.

06

Peer review, publicly reconstructable campaign artefacts, fixed-budget comparisons, and reliable timing remain important next steps.

Reference annex

The analysis above carries the main reading flow. The material below is separated as a reference layer so program teams can inspect terminology, recurring questions, editorial method, and primary sources without interrupting the argument.

Editorial record
Editorial owner
Neura Parse Research
Last verified
July 20, 2026
Method
Synthesis of the dated primary and official records listed below, checked against the operating question in this note.
Scope limit
Planning analysis—not certification, customer performance evidence, procurement advice, or a claim of production readiness.
Choose the next step

Use the beginner roadmap to connect concepts, Qiskit, simulation, optional hardware, research evidence, and a bounded programme decision.