LIBERO#
LIBERO is the default RPent
environment: a MuJoCo/robosuite-based tabletop manipulation benchmark
with four suites (libero_object, libero_goal, libero_spatial,
libero_10) and three variants (standard, pro, plus).
The default VLA is Pi0.5, served over HTTP by
robots/libero/vla_server.py.
VLA configuration#
Pi0.5 needs one thing: a checkpoint on disk. Point at it via
PI05_CHECKPOINT_PATH:
export PI05_CHECKPOINT_PATH=/path/to/rlinf-pi05-libero-130-fullshot-sft
Download the recommended SFT checkpoint from HuggingFace: rlinf-pi05-libero-130-fullshot-sft.
Task selection#
Every LIBERO run picks:
--suite— one of the four suites, each optionally suffixed with the variant flavor (see below). Examples:libero_object_task,libero_object_swap,libero_goal_lan,libero_spatial_task,libero_10_swap.--task— the task index within the suite.--seed— the environment seed.--libero-type— the LIBERO variant:standard|pro|plus. If omitted, RPent falls back toLIBERO_TYPEin the environment (defaultpro).
Suite × variant matrix#
Suite |
Variants |
Purpose |
|---|---|---|
|
|
Object-centric tasks with optional target swap or language perturbations. |
|
|
Goal-conditioned tasks with optional swap / language perturbations. |
|
|
Spatial-relations tasks. |
|
|
The long-horizon LIBERO-10 suite. |
Minimal command#
export PI05_CHECKPOINT_PATH=/path/to/rlinf-pi05-libero-130-fullshot-sft
export LIBERO_TYPE=pro
export CUDA_VISIBLE_DEVICES=0
python rpent/cli/main.py \
--suite libero_object_swap --task 2 --seed 0 \
--cerebrum api --model anthropic:claude-opus-4-8 \
--max-tokens 8192
What runs where#
env_server (
robots/libero/env_server.py) — owns the LIBERO MuJoCo env and EGL rendering. Exposesreset,step,chunk_step,render_agentview,get_camera_meta,cached_image, … over pickle-framed socket RPC.vla_server (
robots/libero/vla_server.py) — owns the Pi0.5 weights. Exposes/predictover HTTP.Toolkit (
robots/libero/toolkit.py) — defines the tools the LLM can call:pi0_pick(fed to Pi0.5),move_to,rotate_wrist,back_project,view_driver_state,finish, …
Tools the planner sees#
By default the LIBERO toolkit exposes:
pi0_pick(target)— invoke Pi0.5 for a pick chunk targetingtarget(a natural-language object description).move_to(dx, dy, dz)— scripted Cartesian motion (deterministic; no VLA).rotate_wrist(delta_rad)— scripted wrist rotation.release()— open the gripper.back_project(pixel_x, pixel_y)— turn a click on the agentview image into a 3D point in world coordinates.view_driver_state()— force a fresh state dump (images, depths, camera meta,states.json).finish(status)— end the episode withsuccess/failure/stuck.
Every tool re-renders the world after it runs, so the next turn’s context reflects the post-action state.
Live dashboard#
Add --dashboard to open a local monitor for a LIBERO run:
python rpent/cli/main.py --dashboard \
--suite libero_goal_task --task 1 --seed 0 --cerebrum claude_code
The dashboard streams reasoning, agentview + wrist camera + Pi0.5
overlays, and an action timeline. Use
--dashboard-language zh-cn for the Chinese UI.
Bringing your own VLA#
If you have a LIBERO-compatible VLA that is not Pi0.5, swap the model client without touching the env by:
Writing a new
vla_server.pythat exposes the same/predictHTTP contract (or a socket-RPC contract, if that fits better).Pointing at it with
--vla-endpoint http://host:port.Optionally updating
robots/libero/toolkit.pyif the tool surface (e.g.pi0_pick→mymodel_pick) needs to change.
See Add an action primitive for the full walkthrough.