Skip to content

Decision Lab · current proof and applied research

Where Physical AI product decisions become inspectable.

Run current decision tools, inspect Auralink and Reachy evidence, and see how product, architecture, model, safety and field assumptions are turned into decision records. The planned physical Foundry remains clearly labelled further down.

An inspectable evidence register linking claims to classifications, sources and decision records.
Current Decision Lab index — evidence remains classified by maturity, source and limitation.

Digital lab live · physical Foundry planned

Runnable digital demonstrations and Hyperion-authored research artefacts are available now. The proposed physical facility, zones and missions remain illustrative until built, tested and labelled otherwise.

Available now

Start with what exists.

These are current digital, bench and internal-R&D artefacts—not the planned physical facility. Each page states what it can and cannot prove.

From Demo to Operated Product · Simulator 1.1

Physical AI Product Flight Simulator

Configure a fictional AMR, drone, cobot or vehicle-feature product. Stress economics, reliability, completeness, recovery and operator burden; inspect the diagnostic gaps; then export Product Decision, Product Bridge and Product System records with provenance. All outputs remain illustrative and experiment-only.

Fly a product decision

Auralink

Auralink

Public Physical AI research · documented by Hyperion. An inspectable research reference—not a client deployment or commercial-outcome claim.

Inspect Auralink

Reachy Mini + SO-101

Reachy Mini + SO-101

Measured · simulated · dry-run. Real sensing and teleoperation evidence stays separate from simulated autonomy and unexecuted robot actions.

Inspect the robotics evidence

Edge-model evaluation

The current field-note revision is awaiting editorial review. Its link and dependent performance figures are withheld; no throughput or product-readiness conclusion is offered.

06 · Available now

Decision tools

Explore six demonstrations and first-pass calculations.

Run the demonstration

Decision Lab protocol · definition only

From observation to operated-product evidence.

Planned · illustrative · no results collected

This protocol defines future records and measures; no run results have been collected for it. The physical Foundry remains planned and illustrative. Apply the protocol at the relevant Product System gate through all six lenses. Future observations belong in the existing Product System records and Evidence Passport, with configuration, exposure, authority, source, limitations, review and expiry. No stage establishes evidence maturity or authorises release automatically.

Five Product System gates and six lenses frame the decision. The five planned Foundry stations describe a facility concept. The six protocol stages below define what to record and measure.

  1. 01

    Observation

    What is true now, how was it observed, and what reference establishes validity?

    Required record

    Timestamped sensor or data observation; origin and rights; configuration and envelope; evaluator or ground truth; missing, invalid and stale-state flags.

    Metric definitions · no measured values

    Valid observation rate

    Valid observations divided by required observation opportunities. Define validity before collection; retain missing observations in the opportunity count.

    Denominator: Required observation opportunities in the declared window.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Observation-to-state latency

    Distribution of elapsed time from a relevant physical event to a usable system state. Declare clock alignment and report censored or missing transitions.

    Denominator: Matched event-to-state pairs; disclose unmatched event count separately.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

  2. 02

    Future

    Which future state is predicted, at what horizon, with what uncertainty and alternatives?

    Required record

    Source observation, forecast horizon, model and configuration, predicted state, uncertainty, alternatives and abstention. Pair predictions with later observations before assessing them.

    Metric definitions · no measured values

    Forecast error by horizon

    Apply a declared loss function between each prediction and its later reference observation. Report a distribution by horizon and slice; unobserved futures remain unassessed.

    Denominator: Matched prediction/reference pairs at each declared horizon; disclose unmatched predictions.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Forecast calibration

    Compare stated probabilities with observed frequencies within predeclared probability bins. Report bin sizes, uncertainty and excluded cases.

    Denominator: Resolved predictions within each declared probability bin and outcome definition.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

  3. 03

    Action

    What was proposed, authorised, executed or rejected, and whose authority governed it?

    Required record

    Typed proposal, applicable constraints, human or deterministic authorisation, bounded command or abstention, execution acknowledgement and actual resulting state. Separate model proposal from control and specialist acceptance.

    Metric definitions · no measured values

    Accepted action transition rate

    Accepted resulting state transitions divided by authorised action attempts. Use predeclared acceptance criteria and retain failed or aborted attempts.

    Denominator: Authorised action attempts within the same configuration and authority boundary.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Authority boundary event rate

    Report blocked, escalated and unauthorised proposals separately by reason, divided by proposal opportunities. A blocked proposal is not inherently a system failure or success.

    Denominator: Proposal opportunities under the declared authority policy.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

  4. 04

    Recovery

    After degradation or failure, was the condition contained and valid task state restored or the run stopped?

    Required record

    Trigger and failed transition, severity, retry/replan/rollback/handback/stop path, recovery owner, restored state, recurrence and outcome. Preserve failed recoveries and human interventions.

    Metric definitions · no measured values

    Accepted recovery rate

    Accepted recoveries divided by recovery attempts, using a declared restored-state criterion. Report degraded, failed, aborted and escalated outcomes separately.

    Denominator: All recovery attempts for the declared failure population and window.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Return-to-duty time

    Distribution of elapsed time from detection to accepted resumed useful work. Report unrecovered and censored episodes separately.

    Denominator: Recovery episodes; disclose which resumed duty and which remained stopped.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

  5. 05

    Useful work

    Did the complete human–machine workflow deliver an accepted customer outcome, with what operator burden?

    Required record

    Customer acceptance definition and human baseline; attempted and accepted outcomes; exposure, scheduled and continuous duty; rejects, aborts, interventions, attention and recovery work.

    Metric definitions · no measured values

    Accepted useful outcome rate

    Accepted useful outcomes divided by attempted outcomes, under the declared acceptance definition. Report throughput and longest continuous useful duty alongside the ratio.

    Denominator: Attempted customer outcomes in the declared window; do not substitute easy subtasks.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Operator burden

    Operator attention, intervention and recovery minutes divided by duty hours. Distinguish elapsed supervision from person-minutes across multiple operators; compare with the human baseline.

    Denominator: Duty hours for the same system, workflow and observation window.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

  6. 06

    Operated-product evidence

    Does this exact configuration sustain value, controls, service and economics for the declared population and window?

    Required record

    Evidence Passport and traceable Product System records: configuration, envelope, authority, population, window, exposure, evidence origin and maturity, source, limitations, review, expiry and human recommendation. Classify gaps; do not promote a planned protocol or simulation into field evidence.

    Metric definitions · no measured values

    Availability by cause

    Available duty time divided by scheduled duty time. Classify unavailable time by cause and disclose planned exclusions and overlapping causes.

    Denominator: Scheduled duty time for the same population and window.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Severity event rate

    Events divided by declared exposure, reported separately for each severity class. Preserve zero events, missing counts and absent exposure as distinct states.

    Denominator: Declared exposure such as attempts, operating hours or distance; specify the unit and never combine incompatible denominators.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

    Cost per successful outcome

    Total ownership and operating cost divided by accepted useful outcomes over the same window. Include integration, infrastructure, service, operator and recovery cost; disclose assumptions and allocation. An absent or zero denominator yields no ratio.

    Denominator: Accepted useful outcomes matching the cost population and accounting window.

    Window and breakdowns

    Declare the observation window, population, configuration, envelope and authority before collecting data. Report missing and excluded cases.

    configuration · operating-envelope slice · authority and supervision · site and task · failure or exclusion reason

These definitions support the existing Product Management metrics tree. Future observations must use the Product System and Evidence Passport. Version: decision-lab-protocol.v1.

The proposition

Not a showroom of tricks. A place to make the next decision.

A model demonstration asks whether one behaviour can work under chosen conditions. The Decision Foundry asks whether the complete product should advance. Customer value, intelligence, hardware, software, safety, operations and economics are evaluated together. A convincing output is not a pass: the evidence determines whether the decision is go, redirect, sequence or stop.

One connected system

One lifecycle. One runtime. One evidence return.

The Product System governs why and when the product advances. The Physical AI system architecture governs how the system behaves in operation.

Product System

Strategy → Discovery → Productization → Launch → Scale

Five gates govern when the product may advance. Six decision lenses test whether the evidence meets the threshold for the current decision.

Physical AI runtime

Sense → Connect → Compute → Reason → Act → Orchestrate

Trust, safety and cybersecurity surround every layer.

Evidence return

Incidents, field variance, technical discovery or economics can reopen an earlier decision. Nothing advances automatically.

The visitor mission

A signal becomes a decision—or stops at the boundary.

01Sense the real decisionOutput: decision contract

Proposed lab zones

  1. 01Mission Lock
  2. 02Field Cell
  3. 03Signal Spine
  4. 04Sovereign Core
  5. 05Knowledge Forge
  6. 06Failure Theatre
  7. 07Authority Gate
  8. 08Fleet Observatory
  9. 09Evidence Chamber
  10. 10Executive Studio
  11. 11Field Notes Studio

Follow one bounded initiative from declared intent to an inspectable decision record. As you move, the architecture resolves around what must be observed, authorized and proved.

  1. 01LOCK

    Sense the real decision

    Name the intended use, accountable owner, system boundary, acceptance threshold and stop condition before a model is selected.

    Output: decision contract
  2. 02GROUND

    Ground and decide

    Observe the declared baseline, connect field signals to governed knowledge and compare options with sources, uncertainty and operating context visible.

    Output: bounded proposal
  3. 03GATE

    Authorize or abstain

    Schema, allow-list, policy and authority checks determine whether a proposal may reach deterministic control. Ambiguity or unsafe intent stops here.

    Output: authorization record
  4. 04RECOVER

    Act, observe and recover

    Run one bounded action, inject declared variance, watch operator and system response, and prove that degradation, stop and rollback paths work.

    Output: observed system behaviour
  5. 05PROVE

    Prove and scale

    Compare the result with acceptance and kill thresholds. Record the evidence, open assumptions and a go, redirect, sequence or stop recommendation.

    Output: Product System gate record

Illustrative architecture

Stress the authority boundary

Change one condition. The model may interpret and propose; the system boundary determines what happens next.

AuthorizedInputs remain inside the declared envelope. A typed proposal passes policy and human authority, reaches deterministic control and returns evidence.

Mistral-only runtime. EU endpoint. Owner-authorized.

Deterministic authority first. Mistral interprets and challenges.

The deployed environment permits deterministic code or exact, pinned Mistral models through a server-only EU endpoint. The production-allowlisted JARVIS, Search and Product Council paths are active by owner decision. Open assurance items remain disclosed; there is no global-region or provider fallback.

  1. Calculate

    Live · deterministic

    Economics, capacity, availability, constraints, risk propagation and hashes are computed by inspectable code with explicit assumptions, units and formulas.

  2. Moderate

    Active · owner-authorized

    The pinned moderation model must approve input and output. If moderation or EU inference is unavailable, AI fails closed while deterministic simulation stays available.

  3. Challenge

    Active · owner-authorized

    Small, Medium and Large serve as separate facilitator, systems-analyst and critic roles. Disagreement is preserved, and shared model-family failure modes are disclosed.

  4. Ground

    Local evidence

    The VPS owns evidence identifiers, retrieval, prompt versions and state. No built-in web search, Files, Libraries, Agents or Conversations service is part of the production path.

  5. Prove

    Versioned evidence

    Every substantive result records exact models, prompt and method versions, classifications, evidence IDs, deterministic hashes, limitations, review status and a human owner.

  6. Contain

    EU-only adapter

    All hosted calls originate server-side at https://api.eu.mistral.ai. Other hosts, model aliases, unapproved IDs and provider fallbacks are rejected; secrets never enter browser code.

No model-to-machine or model-to-decision path

No Mistral response is a direct PLC, robot, charger or vehicle command. Deterministic code owns arithmetic and constraints; an independent safety function may reject or stop action; a named human owns investment, release, compliance, safety and customer commitments.

Explicitly outside production

OCR, embeddings, speech, Agents, Conversations, Files, Libraries, Batch, built-in web search, fine-tuning, Forge and preview services are disabled until the exact capability is verified for EU regional processing, Zero Data Retention, contract fit, security and measured need.

Hyperion Consulting is independent from Mistral AI. The EU regional endpoint is a technical boundary, not proof that the complete subprocessor and control-plane chain remains in the EU. No partnership, certification, sovereignty or compliance status is claimed.

Verify in official Mistral documentation

First mission packs

Designed around failure—not staged perfection.

Each mission begins as planned and illustrative. It becomes measured only when the reference cell exists, the protocol is published and the result is reproducible.

Grounded maintenance

Read machine state and a governed maintenance corpus, show supporting passages, state uncertainty and propose a bounded work order—without autonomously diagnosing equipment or issuing control commands.

Planned · illustrative until commissioned

Authority boundary

Give the system an ambiguous or unsafe instruction. Intelligence may propose or abstain; typed interfaces, policy authority and independent safety determine whether anything can reach the controller.

Planned · illustrative until commissioned

Pull the network

Disconnect external connectivity and observe local inference, cached knowledge and deterministic fallback against declared degraded-mode requirements.

Planned · illustrative until commissioned

Operating envelope

Challenge perception and learned behaviour with glare, occlusion, unfamiliar parts, drift and workstation variation to expose—not hide—the validated boundary.

Planned · illustrative until commissioned

Canary to fleet

Test a revision in simulation, release it to one controlled asset, expand only while acceptance thresholds hold and roll back when they do not.

Planned · illustrative until commissioned

Evidence Passport

Every impressive moment carries its limits.

A demonstration is useful only when a buyer can tell what happened, under which conditions and what the result does not prove.

Live
Runnable now; not automatically production-proven.
Measured
Observed under named conditions; not extrapolated beyond them.
Simulated
Produced in a digital or synthetic environment; not field operation.
Replay
A fixed prior run; not a live interaction.
Illustrative
A proposed scenario or architecture; implementation is not implied.
Planned
Intended for a future build; not commissioned.

The record travels with the result

Question · intended use · data origin and rights · model, prompt, tool, firmware and dataset revisions · operating conditions · authority boundary · acceptance and kill thresholds · expected and observed result · failures · limitations · reproduction date · what the run does not prove

JARVIS · Lab OS

The guide—not the governor.

JARVIS lets a visitor observe system state, stress a bounded scenario, compare decisions and explain the result against cited evidence. It may synthesize and propose; it cannot waive an acceptance threshold, certify a system or authorize unsafe physical action.

Ask JARVIS about the Foundry
Observe
System state, provenance and sources
Stress
Bounded failure injection
Decide
Options and lifecycle consequences
Explain
Evidence, limitations and decision record

How the complete system works

One Physical AI system — from sensing to dependable action

Select a stage to follow the system end to end.

Pixel-art diagram of the Physical AI stack, with signal rising from sensing through orchestration.

01 / 08Device & Control

Sensors & environment. The physical world, sensed — cameras, depth, radar and OT signals.

See how we engineer this

Decision Foundry questions

Is the Decision Foundry already open?

No. The physical facility is in concept development and is not commissioned. Current digital demonstrations and Hyperion-owned R&D are labelled separately.

Is Hyperion a Mistral AI partner?

No partnership or special access is claimed. Hyperion is independent from Mistral AI. The owner-authorized JARVIS, Search and Product Council paths use a Mistral-only runtime through the EU endpoint, with deterministic fallbacks.

Can a Mistral model directly control a machine in the proposed lab?

No. Model outputs are proposals that must cross schema, allow-list, policy and authority checks before deterministic control. Independent safety remains outside model authority.

Is the Decision Foundry a certification laboratory?

No. It does not replace legal, safety, cybersecurity, conformity-assessment or certification specialists. It creates product and architecture evidence for a bounded decision.

What should a participant leave with?

A bounded decision record: observed evidence, unresolved assumptions, limitations and a go, redirect, sequence or stop recommendation.

Which industries does the programme cover?

Manufacturing, automotive and energy are primary contexts. Smart infrastructure, logistics and defence are explicitly labelled research contexts; publication does not imply sector client delivery.

Bring one stuck initiative

Leave with the next decision.

In a 30-minute fit call, Mohammed will identify the decision, the missing evidence and whether a Foundry mission or one of Hyperion’s three mandates is appropriate—or say plainly that it is not.

30 minutes · no obligation · Mohammed leads every conversation