rskill-playbook-verify_outcome

A kind: playbook rSkill: a symbolic S2 decision procedure the Reasoner reads, not a neural policy. It carries no weights β€” the authored PLAYBOOK.md is its runtime.

What this skill does

Closes the loop after a skill or subtask completes (Inner Monologue). The action result ("the policy reported done") is not the same as the world state, so this playbook verifies what actually happened: it asks the scene a specific yes/no success question, cross-checks a reward monitor, optionally confirms the object's pose, and classifies success or failure with evidence before the reasoner proceeds. A failure routes to a replan or human handoff β€” never a silent "success". Concrete walkthrough: the black-bowl-on-plate example in PLAYBOOK.md.

How it works

This playbook is content, not code. When installed, the reasoner injects PLAYBOOK.md into its system prompt and follows the SOP, composing tools it already has (query_scene, query_task_progress, locate_in_view, memory_write, emit_prompt). It is role: s2 and is never dispatched through ExecuteSkill. Any motion a downstream replan triggers is an execute_rskill β†’ Action chunk β†’ C++ safety kernel β€” the playbook holds no actuation authority (CLAUDE.md Β§1.1).

Observation β†’ action contract

None. A playbook emits no Action chunks and requires no actuators (actuators_required: [], chunk_size: 1). Its "output" is the sequence of tool calls the reasoner makes while following the SOP, bounded by playbook.max_steps.

How it was authored / Upstream provenance

N/A β€” a playbook is hand-authored, not trained: it has no weights and no upstream model. Its provenance is the authoring decision record (also linked via paper_url). To change behaviour, edit PLAYBOOK.md and bump version.

Supported robots

Embodiment-agnostic β€” declares the explicit wildcard embodiment_tags: ["any"] (never an empty list). Gated by capabilities_required (has_vision: true β€” a real RobotCapabilities flag): the loader filters it out on robots without a camera. The scene / reward queries are gated at runtime by the composed tools, not by this playbook's flags.

Sensors required

None directly. The tools it composes declare their own sensor needs.

Manifest summary

  • kind: playbook, role: s2, actions: [plan], chunk_size: 1.
  • playbook.trigger: a skill or subtask has just completed and its success is not directly observable from the action result.
  • playbook.done_predicate: the outcome is classified success or failure with evidence, and a failure has triggered a replan or handoff.
  • playbook.max_steps: 6.

Quick start

from openral_core.schemas import RSkillManifest

m = RSkillManifest.from_yaml("rskills/verify-outcome/rskill.yaml")
assert m.kind == "playbook" and m.playbook is not None
print(m.playbook.trigger)

Reproduction

Packaging-only: the manifest + SOP are validated by tests/unit/test_playbook_rskill_manifest.py. There is no benchmark number to reproduce; the playbook's behaviour is exercised by the reasoner integration tests in later phases.

Evaluation

N/A β€” no eval/*.json; a playbook produces no benchmarkable policy output.

License

  • Code / content: Apache-2.0.
  • Weights: none.

See also

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading