Agent OS & Skill Agent Framework
1. Skill Agent Creation via Markdown (SKILL.md)
Skill agents are defined entirely through Markdown files — no code required. Each skill is a SKILL.md file that describes the agent's persona, instructions, available tools, and behavior rules in natural language.
How It Works
- Create or edit a
SKILL.mdfile describing the skill's purpose, persona, tool list, and constraints - The framework reads and parses the markdown at inference time
- The parsed content becomes the system prompt for the LLM during the ReAct execution loop
- No redeployment or restart needed — changes to
SKILL.mdtake effect on the next inference call
Agent Directory Layout
Each agent is organized under its department with a consistent folder structure:
agent_workspaces/{department}/agentos_agents/{agent_id}/
├── agent_config.json # Agent metadata (name, model, department)
├── enterprise_context/
│ ├── Enterprise_Context.md # Org-wide context injected into all agents
│ └── policies/*.md # Compliance & policy documents
├── skills/
│ ├── _index.yaml # Skill registry (routes queries to skills)
│ └── {skill_name}/
│ ├── SKILL.md # Skill definition (persona + tools + rules)
│ ├── config.json # Skill-level config (model overrides, params)
│ └── additional_files/ # Supporting files (CSVs, templates, etc.)
└── .audit/
└── tool_audit_*.jsonl # Audit trail of tool executions
2. ReAct Execution Loop
The framework uses a single, robust execution mode: ReAct (Reason + Act). The LLM iteratively reasons about the user's query, selects a tool, observes the result, and repeats until it can provide a final answer.
Loop Mechanics
- Reason — LLM analyzes the current state (user query + tool results so far)
- Act — LLM selects a tool and provides arguments
- Observe — Tool executes, result is appended to conversation
- Repeat — Loop continues until LLM produces a final answer (no tool call)
Key Behaviors
- Max iterations guard prevents infinite loops
- Human-in-the-loop integration pauses execution mid-loop when approval is required
- Token budget tracking — each iteration's token usage is accumulated
- Error resilience — tool failures are reported back to the LLM as observations, allowing graceful recovery
3. Skill Routing (Multi-Skill Agents)
When an agent has multiple skills, the framework automatically routes the user's query to the most relevant skill using an LLM-powered router.
How Routing Works
- A skill index (
_index.yaml) lists all available skills with short descriptions - The router presents the skill list to the LLM along with the user's query
- The LLM selects the best-matching skill
- The selected skill's
SKILL.mdis loaded and used for the ReAct loop
Fallback Behavior
- If only one skill exists, routing is skipped — the single skill is used directly
- If the router cannot determine a match, it defaults to the first skill or asks the user for clarification
4. Human-in-the-Loop (HITL)
The framework supports approval gates where sensitive tool executions pause and wait for human approval before proceeding.
How HITL Works
- A tool is marked as requiring approval (via hook configuration or skill config)
- During the ReAct loop, when the LLM requests that tool, execution pauses
- The pending tool call (name + arguments) is returned to the UI
- The user reviews and approves or rejects the tool call
- On approval, the loop resumes from where it paused; on rejection, the LLM is informed and re-plans
Integration
- Works with the Hook System (see §8) — PreToolUse hooks can trigger approval gates
- UI surfaces the pending approval with full tool name and arguments for transparency
5. Enterprise Context Injection
Every agent automatically receives organization-wide context that shapes its behavior, policies, and domain knowledge.
Context Sources
| Source | Purpose |
|---|---|
Enterprise Context (.md) |
Org-wide rules, domain knowledge, terminology |
Policy Documents (.md) |
Compliance policies, SOPs, security guidelines |
How It Works
- At inference time, the framework reads all enterprise context and policy files
- This content is prepended to the system prompt (before the skill-specific content)
- The LLM sees enterprise context as authoritative background — it cannot be overridden by user queries
- Enables consistent behavior across all agents in a department
Use Cases
- Enforce compliance rules across all agents
- Inject company-specific terminology and acronyms
- Provide org chart or escalation paths
- Set tone/formality standards
6. Multi-Skill Composition
A single agent can host multiple skills, each defined by its own SKILL.md. This allows building agents that cover broad domains without creating separate agents for each sub-task.
Example
Agent: "IT Operations Assistant"
├── Skill: network_troubleshooter → Diagnoses connectivity, DNS, firewall issues
├── Skill: server_provisioning → Provisions VMs, configures OS, sets up monitoring
├── Skill: incident_management → Creates/updates tickets, escalates incidents
└── Skill: knowledge_base_search → Searches internal docs and runbooks
Benefits
- Single entry point — users interact with one agent regardless of task type
- Separation of concerns — each skill has its own tools, rules, and persona
- Independent updates — modify one skill without affecting others
- Automatic routing — the framework picks the right skill based on the query
7. Token Usage & Cost Tracking
The framework tracks token consumption per inference call, providing visibility into LLM costs.
Tracked Metrics
- Prompt tokens — tokens sent to the LLM (system prompt + conversation + tool results)
- Completion tokens — tokens generated by the LLM (reasoning + tool calls + final answer)
- Total tokens — sum of prompt + completion
- Per-iteration breakdown — token usage for each ReAct loop iteration
Benefits
- Token counts are returned in the inference response payload
- Can be aggregated for billing, budgeting, and optimization
- Helps identify skills that are token-heavy and need optimization
8. Hook System (Pre-Tool, Post-Tool, Pre-Response)
The framework provides a powerful hook system that intercepts execution at key points, enabling custom logic without modifying core agent code.
Hook Types
8.1 PreToolUse Hook
Fires: Before a tool executes
Can: Modify tool arguments, block tool execution, require human approval
Use Cases:
- Block dangerous commands (e.g.,
rm -rf /) - Sanitize inputs (e.g., redact PII from queries)
- Enforce approval for production-impacting tools
- Rate-limit tool calls
8.2 PostToolUse Hook
Fires: After a tool executes, before the result reaches the LLM
Can: Modify tool output, block result propagation, log/audit
Use Cases:
- Redact sensitive data from tool outputs before LLM sees it
- Enrich tool results with additional context
- Log all tool outputs for compliance
- Transform output format
8.3 PreResponse Hook
Fires: After the LLM produces its final answer, before sending to the user
Can: Modify the response, block it, add disclaimers
Use Cases:
- Append compliance disclaimers
- Filter inappropriate content
- Add citations or source links
- Format response for specific channels (Teams, Slack, etc.)
External Script Hooks
In addition to built-in hooks, the framework supports external hook scripts that run before/after tool execution. This allows teams to plug in custom validation logic written in any language.
Configuration
Hooks are configured per-agent or globally. They can be enabled/disabled without restarting the server.
Summary
| # | Feature | Description |
|---|---|---|
| 1 | Skill Creation via SKILL.md | Define agents using natural language markdown — no code |
| 2 | ReAct Execution Loop | Iterative reason-act-observe loop for reliable task completion |
| 3 | Skill Routing | LLM-powered query routing across multiple skills |
| 4 | Human-in-the-Loop | Approval gates for sensitive tool executions |
| 5 | Enterprise Context | Org-wide policies and context injected into all agents |
| 6 | Multi-Skill Composition | Multiple skills per agent with automatic routing |
| 7 | Token Tracking | Per-call token usage metrics for cost visibility |
| 8 | Hook System | Pre-tool, post-tool, and pre-response interception points |