复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Autonomous OS is the open source operating system for physical AI agents.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Autonomous OS is the open source operating system for physical AI agents. It runs on edge devices with cameras, microphones, speakers, displays, motors, lights, and sensors, and gives an AI agent a body: it sees, hears, speaks, moves, senses, remembers, runs skills, and updates itself — locally first.
https://github.com/user-attachments/assets/08a8f109-a165-4604-be5f-39e5ec9e8d11
Autonomous Lamp is the first reference device. Intern is the second. Anyone can build a third.
The brain is a swappable agentic runtime (OpenClaw, Hermes, OpenCode, OpenAI Codex, Claude Code, or any LLM + skills + memory). Autonomous OS is everything else — the body, the skills, and the bounds.
| Device | What it is | Declares | |
|---|---|---|---|
| Autonomous Lamp | 5-DOF expressive desk robot | the maximal set — audio, vision, motion, light, display, sensing | |
![]() | Autonomous Intern | always-on desk agent | audio, sensing, light — no camera, motion, or display |
| Reachy Mini | Pollen Robotics' desk robot, running Autonomous | audio, vision, motion (6-DOF head + 360° body), expression, sensing — no light or display | |
![]() | Unitree Go2-W | a different manufacturer's mobile robot, running Autonomous | audio, vision (+ depth), motion (locomotion), sensing |
Lamp and Intern are Autonomous's own devices; Reachy Mini and the Unitree Go2-W belong to
other manufacturers and run the identical OS — the Android playbook (Android on Samsung,
Pixel, …). They all run the same OS image; only their DEVICE.md differs.
Reachy Mini is where that stops being a claim. It is someone else's hardware, shipped with
its own vendor OS and daemon, and Autonomous installs alongside it —
one command on the robot. Onboarding it was writing a DEVICE.md, a
motion driver wrapping Pollen's Python SDK, and a SAFETY.md. No fork. Its motion is a
Stewart platform — 6-DOF parallel kinematics, nothing like Lamp's serial bus servos or the
Go2-W's legs — yet a skill calling motion.aim runs on all of them, because skills address
capabilities, never hardware.
Autonomous OS is a layered stack: each layer exposes an interface to the one above and depends only on the one below, so any layer can be replaced without touching the others.
What the device does — 24 skills, each a SKILL.md the runtime invokes: apps like guard,
mood, scene, habit, wellbeing, plus capability wrappers (led-control,
servo-control, camera, music, …). A skill is an ability; the device's character is
its SOUL.md. First-party skills use the same public contract a third party gets. (skills/)
The always-on Go daemon: intent (fast local commands), network, sensing routing,
monitor (flow event bus), healthwatch, ambient, and device. Deterministic — they run
with or without the runtime. OTA runs as its own worker (bootstrap/).
(system/)
OpenClaw, Hermes, OpenCode, OpenAI Codex, Claude Code, or a custom
runtime. Runs the skills, embodies the device's SOUL.md, and decides what to act on.
Swappable at runtime (web Settings or MQTT) — and where Autonomous OS's differentiated value
(the default brain, memory, character) lives. Its tools — how it reaches beyond the device —
are MCP connectors (runtimes/*/mcp.go, synced across a switch by system/agent) and the
CLI the LLM calls directly (curl, shell); skills are the device's own abilities through the
HAL, tools are external capabilities the runtime calls.
(runtimes/{openclaw,hermes,opencode,codex,claudecode}; adding
your own: docs/agentic/adding-agent-runtime.md)
The frozen, versioned interface — 12 capabilities: audio, vision, sensing, presence,
motion, light, display, expression, media, connectivity, companion, system.
Skills call capabilities (motion.move), never hardware models — so one skill runs on any
body that declares the capability. A device's DEVICE.md declares which it has; the runtime
mounts only those. The HAL also hosts the safety gate (hal/safety): SAFETY.md
bounds — e-stop, motion limits, brightness, quiet hours — enforced deterministically below
the brain, never by the LLM.
(devices/contract/ + hal — see HAL)
The realtime voice agent (hal/realtime) — brain-tier code the HAL hosts in-process, so it sits
between the runtime and the HAL. Voice turns land here first and it decides per turn: answer
directly when the turn is simple (small talk, nothing that needs skills or tools), or delegate
up to the main agentic runtime when the turn needs skills or complex tool calls. Runs on Gemini
Live, OpenAI Realtime, or Qwen.
(hal/realtime — see realtime-voice.md)
The vendor kernel (Raspberry Pi OS / OrangePi, or the robot's onboard compute) we run on — we
don't ship one. Our Drivers (motors, rgb, display, camera, voice (STT/TTS/VAD),
gpio/touch, bluetooth in hal/drivers, with per-board wiring in hal/board) are
userspace programs talking to it through GPIO/SPI/ALSA/V4L2;
Power Management is the foundation.
(see kernel)
📖 Full docs: overview · HAL · kernel
Every device is self-describing to both humans and the runtime, in four files:
| File | Role | Consumer |
|---|---|---|
DEVICE.md | the body — what hardware is present | the OS, at boot |
SKILL.md | the hands — what it can do | the runtime |
SOUL.md | the self — who it is | the runtime |
SAFETY.md | the bounds — what it must never do | the OS (deterministic) |
The contract that governs them lives under devices/contract/ — see
DEVICE-SPEC.md and capabilities.md.
The tree maps onto the architecture layers (top of the stack first):
# The OS
skills/ Skills — the apps (SKILL.md)
system/ System Managers (Go): one folder per manager — intent, network, monitor, OTA…
web/ on-device setup + monitor UI (React)
runtimes/ Agentic Runtime — one folder per swappable brain (openclaw, hermes, opencode, codex, claudecode)
hal/ HAL (Python) — the package; capability host + routes
drivers/ Drivers — by subsystem (motion, audio, vision, light, display, sensing)
board/ Board Support — per-board profiles + declaration-driven mounting
devices/ reference devices: lamp/, intern-v2/, reachy-mini/, unitree-go2w/ (DEVICE · SOUL · SAFETY · README)
contract/ HAL capability ABI — frozen, versioned (what skills build against)
cts/ compliance test suite — validates devices against the contract
# Supporting
docs/ documentation, incl. docs/architecture/
scripts/ build, OTA, and SBC image tooling (incl. scripts/imager/)
# Off-device & integrations
integrations/
companions/ desktop companion apps (autonomous-buddy, claude-desktop-buddy)
chat-bridges/ chat bridges into the device (Twitch, web chat)
perception-service/ off-device cloud perception inference
DriversandBoard Supportare surfaced ashal/driversandhal/board.
# Go system services (cross-compiled to linux/arm64 — Pi or OrangePi)
make os-build # builds the system server (system/)
make os-test # go test ./...
# Hardware runtime (runs on the Pi or OrangePi)
cd hal && uv sync
make hal-dev # uvicorn reload on :5001
make hal-test # pytest
# Web UI
make web-install && make web-dev
All HTTP endpoints return {"status": 1, "data": <payload>, "message": null} on success
and {"status": 0, "data": null, "message": "error"} on failure.
Apache 2.0 — fully open. Build a device by writing a DEVICE.md, a driver, and a
SOUL.md; you never fork the OS. PRs welcome — vibe-coded ones too 🤖. See
CONTRIBUTING.md.
name: mood
description: Tracks the USER's mood only — signals + synthesized decision from camera/voice/telegram. Do NOT use for emotion commands directed at the device ("show sad", "be happy", bare "sad now"); those go through emotion/SKILL.md and are never logged here. Music/wellbeing skills consume the latest decision.OUTPUT RULE — read this before you type anything to the user.
This skill is an internal workflow. NEVER narrate it into your reply. Forbidden in the reply text:
- Section names or step numbers ("Step 1", "Workflow", "After Logging Decision", "Flow A").
- Phrases like "Now I follow…", "Let me check…", "Next step…", "I'll log…".
- Bullet lists re-hashing the mood history you just read ("- Normal (15:00) — …" / "- Excited (16:00) — …").
- The mood value itself as a label ("Mood: sad", "Decision: happy").
- Any of the JSON / curl / timestamps from this skill.
Your reply text to the user is at most ONE short caring sentence (or
NO_REPLY). All the workflow, logging, and synthesis happen silently via tool calls — the user only hears what you would naturally say if you were truly noticing how they feel.
ALWAYS log.
unknownis a validuservalue — log signals and decisions underuser: "unknown"whencurrent_useris unknown. Never skip logging because the user is unknown/unconfirmed; stranger mood still counts for Music decisions.
Mood is stored as two kinds of rows:
signal — raw evidence from one source (camera action, voice tone, telegram message). Multiple per minute is fine.decision — your synthesized mood after looking at the recent signals + the previous decision. This is the row downstream skills (Music, Wellbeing) read.You are the synthesis. The store does not fuse anything. Every time a signal comes in, you log it raw, then immediately read recent history and append a fresh decision row.
happy, sad, stressed, tired, excited, bored, frustrated, energetic, affectionate, unwell, normal
normal is the baseline when nothing strong is going on (use it for decisions when signals are sparse or stale).
| Source | Examples |
|---|---|
camera | facial action: laughing, crying, yawning, sneezing, hugging, kissing, headbanging |
voice | tone: soft, raised, sigh, laugh, monotone |
telegram | message text: "lots of bugs today", "I'm tired", "let's gooo" |
conversation | inferred from a stretch of voice/chat over multiple turns |
| Action | Mood |
|---|---|
| laughing, singing | happy |
| crying | sad |
| yawning | tired |
| applauding, clapping, celebrating | excited |
| sneezing | unwell |
| hugging, kissing | affectionate |
| headbanging | energetic |
For voice/telegram, infer boldly from a single line ("work is killing me" → stressed). Trust your gut.
Skip only if: quoting someone else, or speaking purely hypothetically.
When this skill runs as part of the emotion pipeline — either emotion.detected (camera) or speech_emotion.detected (voice) — the backend injects an [emotion_context: {...}] block with everything you need pre-computed:
recent_signals — array of {age_min, mood, source, trigger} for signals within the last 30 minutes.prior_decision — the most recent kind=decision row as {mood, age_min}, or null.is_decision_stale — boolean (age_min >= 30 or no decision today).Do NOT GET mood-history again in that case — use the context block.
When the skill runs from another path (voice/telegram-driven mood signal, no [emotion_context:] block), fall back to:
curl -s "http://127.0.0.1:5000/api/openclaw/mood-history?user=<name>&last=15"
This returns the full ordered list {signal, decision}; derive the same three fields locally. The GET should batch concurrently with any other reads in the same turn (no data dependency).
Apply this judgment when synthesizing the fused mood:
normal.happy but telegram says stressed in the same window → trust the higher-bandwidth source. Words about feelings beat a momentary facial expression. Multiple aligned signals beat a single outlier.tired after a stressed decision) → shift, don't snap.Embed both rows at the start of your spoken reply as HW markers. The runtime parses them, fires the POSTs in parallel goroutines, and strips them before TTS speaks the rest.
Signal row (raw evidence):
[HW:/mood/log:{"kind":"signal","mood":"<mood>","source":"<camera|voice|telegram|conversation>","trigger":"<short reason>","user":"<name>"}]
Decision row (synthesized):
[HW:/mood/log:{"kind":"decision","mood":"<fused mood>","based_on":"<short summary>","reasoning":"<why>","user":"<name>"}]
Both markers can sit in the same reply (signal first, then decision is fine — they fire concurrently anyway). They use the same endpoint; kind in the body distinguishes them.
| Field | Required | Notes |
|---|---|---|
kind | Yes | signal or decision |
mood | Yes | from the values list above |
based_on | Decision only | e.g. "3 signals last 20min + last decision (stressed, 18min ago)" |
reasoning | Decision only | one sentence, e.g. "telegram complaints outweigh the smile from camera" |
user | No | omit to use current presence user |
Do NOT use curl exec for these logs. Each curl consumes a tool turn (~5-7s LLM-think on the result) for a side-effect with nothing to wait on. The HW marker path is single-trip.
Regex caveat: the marker body must not contain }. based_on / reasoning are usually plain English so this is rarely a problem; if a value would contain } use the curl fallback instead.
curl -s -X POST http://127.0.0.1:5000/api/mood/log \
-H 'Content-Type: application/json' \
-d '{"kind":"signal","mood":"<mood>","source":"...","trigger":"...","user":"<name>"}'
curl -s -X POST http://127.0.0.1:5000/api/mood/log \
-H 'Content-Type: application/json' \
-d '{"kind":"decision","mood":"<fused>","based_on":"...","reasoning":"...","user":"<name>"}'
source is automatically set to "agent" for decisions; do not pass source or trigger.
user — face recognition sets the current user. If you need to verify, query GET http://127.0.0.1:5001/face/current-user → {"current_user": "<name>"} (friend name, "unknown" for strangers-only, or empty string when nobody is present). Do NOT parse this out of /face/cooldowns — that endpoint is for the friend/stranger cooldown debug view, not for attribution.[telegram:SenderName], lowercase.unknown).unknown users too — Music still suggests for them.On emotion.detected and speech_emotion.detected turns, user-emotion-detection/SKILL.md is the router — it picks one of music / checkin / action / silent and gates whether music-suggestion/SKILL.md fires this turn. Voice and camera share one cooldown and one decision row schema; the only thing that changes per modality is the source field on the raw signal row.
When the router picks music (decision mood is suggestion-worthy — sad, stressed, tired, excited, happy, bored — and audio is idle, cooldown clear, decision fresh), the decision POST and the music-suggestion POST share a single write batch — do not split them across tool turns.
Other moods (frustrated, energetic, affectionate, unwell, normal) take a non-music route (checkin / action / silent per the router table) and skip the music POST.
For unknown users — still suggest (speak only, no DM) on the music route. See music-suggestion/SKILL.md for details.
Camera detects yawn, no recent context:
{"kind":"signal","mood":"tired","source":"camera","trigger":"yawning"}{"kind":"decision","mood":"tired","based_on":"1 fresh signal, no recent decision","reasoning":"single yawning signal after stale window"}tired from the same turn (suggestion-worthy).Telegram says "let's go!" but camera 5 min earlier said yawning:
tired (camera, 5min ago), excited (telegram, just now). Last decision: tired, 4min ago.{"kind":"signal","mood":"excited","source":"telegram","trigger":"let's go!"}{"kind":"decision","mood":"excited","based_on":"telegram excitement overrides 5min-old camera yawn","reasoning":"verbal enthusiasm is higher-signal than a single facial cue"}excited.Quiet evening, no recent signals, user just sat down:
normal after the next signal arrives.
评论 (0)
暂无评论,成为第一个评论者吧!