BUSINESS

Language Models Don't Understand Physics. That's Half Your Use Cases.

Published September 27, 2026 — 4 min read

TL;DR: A February 2026 survey of LLM reasoning failures (Song, Han & Goodman) formally separates "embodied" reasoning — the physical, spatial, causal kind — from the language-pattern reasoning models are actually trained on, and finds the embodied category is where failures are most persistent. If your use case touches a warehouse floor, a production line, or a delivery route, you're deploying a language engine into a physics problem, and that mismatch is why manufacturing and logistics AI projects miss their numbers more often than back-office ones.

Key Insight

The industry pitch for LLMs is "general intelligence" — one model, any task. The reality is narrower: LLMs are extremely good at next-token prediction over text, which makes them excellent at anything that reduces to language, code, or symbolic pattern-matching, and unreliable at anything that reduces to physical-world state.

The arXiv survey "Large Language Model Reasoning Failures" makes this precise instead of anecdotal. It splits reasoning into embodied and non-embodied types, and further splits non-embodied into informal (intuitive) and formal (logical) reasoning. The point isn't that LLMs fail randomly — it's that failures cluster by category, and embodied reasoning (spatial relationships, physical affordances, real-time state tracking, cause-and-effect in a physical system) is a distinct, harder category that text-trained models were never optimized for. A model can ace a bar exam and still be unable to reliably reason about whether a pallet fits through a doorway.

This is the framework executives are missing when they greenlight AI in physical operations: the question isn't "is this model smart enough," it's "is this task language-shaped or physics-shaped."

Why Teams Miss This

The mistake is treating model capability as a single scalar — GPT-5-class vs. GPT-4-class, bigger vs. smaller — instead of asking what kind of reasoning the task requires.

How to Actually Do It

Sort your AI use cases by reasoning type before you sort them by model tier:

  1. Classify the task first. Ask: does this task resolve by manipulating language/symbols (drafting, summarizing, classifying, code generation, querying structured data) or by reasoning about physical state (spatial layout, motion, real-time sensor data, mechanical cause-and-effect)? The first category is squarely in LLM strength; the second needs help.

2. Pair the LLM with a system that actually models the physical world. For embodied tasks, don't ask the LLM to reason about physics directly — use it as the language/orchestration layer on top of purpose-built systems: computer vision for spatial state, simulation or digital-twin models for physical prediction, control systems for real-time actuation. The LLM's job becomes translating between human intent and those systems, not doing the physical reasoning itself.

3. Constrain physical-reasoning outputs with verification, not just prompting. If an LLM is going to output anything that touches a physical system (a route, a placement instruction, a torque setting), route that output through a deterministic validator — a physics check, a constraint solver, a simulation replay — before it reaches hardware or a human executor. Treat the LLM's physical suggestions as drafts, not conclusions.

4. Re-audit your existing physical-ops AI projects against this split. For any manufacturing/logistics AI initiative already underway, identify which specific decisions the model is making unsupervised. If those decisions are embodied-reasoning calls (not language calls dressed up as physical ones), that's your highest-risk failure point — not model choice, not prompt engineering.

# Rough shape: LLM as orchestrator, not physical reasoner
def handle_warehouse_task(request):
    intent = llm.parse_intent(request)          # language reasoning — LLM's strength
    plan = digital_twin.simulate(intent)          # physical reasoning — purpose-built system
    if not physics_validator.check(plan):         # deterministic verification
        return llm.explain_infeasibility(plan)   # back to language reasoning
    return plan

What We've Learned

Before approving the next AI pilot in a physical-operations domain, run this filter: list the five decisions the model will make, and mark each one language-shaped or physics-shaped. If more than half are physics-shaped, the feasibility question isn't "which model" — it's "what's the non-LLM system doing the physical reasoning, and is the LLM just the interface to it." That reframe changes the vendor questions, the pilot scope, and the failure mode you're actually testing for.

Sources

Have a specific workflow in mind?

Bring it to a Quick Scan — a live working session where we'll tell you honestly whether it should be an agent, a workflow, or left alone, before you spend a dollar building it. You get 3 prioritized recommendations on the call, a one-page summary after, and the $500 credited toward any engagement within 30 days.

See AI Agent Consulting →·Book an intro call →

Get new posts + practical agent-ops notes

One email when something new goes up. No nurture sequence, no spam — unsubscribe whenever you want.

Thanks — you're on the list.