StrategyCore
Back to Resources
36/ai / ml 8 min read

Human Judgment as a Design Requirement

Ver1.2 of the AI Guidelines for Business moved human oversight from something you promise to something you build. This is what changes in the product when a review asks you to prove it.

Most products describe human oversight as a customer responsibility. The buyer may review outputs, escalate anything unusual, and configure approvals to taste. That framing worked while AI adoption in Japan sat in pilots. It reads badly now, because the guidance a reviewer holds treats human judgment as part of the design, and a capability the customer has to assemble themselves is not a design.

01

What actually changed in March 2026

METI and MIC issued Ver1.2 of the AI Guidelines for Business on 31 March 2026. The revision reflects a shift from experimentation to social implementation, and it defines AI agents and physical AI for the first time, which matters because both act on the world rather than merely describing it. Alongside those definitions sits a design philosophy that makes human judgment essential. The guidance is soft law and carries no penalty. It is also the document your buyer's governance function reads, so deviation is permitted and unexplained deviation is not.

02

A checkpoint is a feature, not a setting

The distinction a reviewer is drawing is between a system where a person must act for a decision to complete, and one where a person may look if they choose to. The first is a control. The second is a hope with a log file. Building the first means an action that halts, a queue it halts into, a role authorised to release it, and a record of who released it and when. That is a week of product work in most systems, and it converts an entire category of review objection into a screenshot.

03

Choose the decision points deliberately

Putting a person in front of everything destroys the economics and trains reviewers to approve without reading, which is worse than no review. The useful question is where a wrong answer is expensive and hard to reverse. Money leaving the company. A message going to a customer. A record being deleted. A permission being granted. Those deserve a halt. Summarising a document for internal reading does not, and claiming it does will make a reviewer wonder whether you understand your own risk profile.

04

Agents raise the stakes on the same question

Ver1.2 naming AI agents is a signal worth taking seriously. An agent that calls tools is not producing text for a person to evaluate; it is taking actions whose effects exist whether or not anyone reads the transcript. The oversight question becomes narrower and harder. Which tool calls execute without a person, what permissions does the executor hold, and what stops a chain of individually reasonable steps from arriving somewhere nobody sanctioned. Products that can answer that at the tool boundary are in a different conversation from products that answer it at the prompt.

05

Show it rather than describe it

The evidence a reviewer wants is mundane. A screenshot of the approval queue. The permission model for who can release an item. An audit record showing a real decision, with the reviewer identity and timestamp intact. A note on what happens when nobody approves within the window, because a queue that silently auto-releases at timeout is the control failing quietly. Assemble that once and it serves every subsequent review, in Japan and increasingly outside it.

// Key Takeaways

What to remember

  • Ver1.2, issued 31 March 2026, treats human judgment as part of the design rather than a customer configuration task.
  • A control is an action that cannot complete without a person. Anything weaker is a hope with a log file.
  • Put checkpoints where a wrong answer is expensive and hard to reverse, and nowhere else.
  • For agents the question moves to the tool boundary: which calls execute unattended, and under whose permissions.
  • A queue that auto-releases on timeout is the control failing quietly, and reviewers look for exactly that.
  • One folder of evidence, assembled once, serves every review that follows.

Last updated:

Scoping Japan entry in this category?

If your company is weighing Japan entry in the work above, StrategyCore is the operating layer that carries it from first assessment to live deployments, run locally and in Japanese.