10 KiB
Multi-Model Skill Pipelines
Overview
Didactyl's skill system already supports executing a series of tasks where each task uses a different LLM model — even from different providers. This document describes what works today, what the limitations are, and what improvements would unlock more powerful multi-step workflows.
What Works Today
Per-Skill LLM Override
Every skill can declare its own llm tag with a provider/model spec and fallback chain:
["llm", "anthropic/claude-sonnet-4-20250514, cheap"]
["llm", "openai/gpt-4o-mini"]
["llm", "best"]
When a triggered skill fires, the runtime applies the skill's execution parameters (model, temperature, max_tokens) before calling the LLM, then restores the agent defaults afterward. This happens in apply_trigger_runtime_to_llm_config() in src/trigger_manager.c.
The llm tag supports the provider/model format. If the tag contains a slash, the runtime parses the provider name and model name separately and overrides both in the LLM config for that execution.
Chain Triggers Connect Skills Sequentially
The chain trigger type fires when another skill completes execution. The filter field specifies the source skill's d tag:
{
"trigger": "chain",
"filter": "source-skill-d-tag"
}
After a triggered skill completes, trigger_manager_fire_chains() in src/trigger_manager.c looks for all adopted skills with trigger=chain whose filter matches the completed skill's d-tag, and fires them.
Combined: Multi-Model Pipeline
By combining per-skill LLM overrides with chain triggers, you can build a pipeline where each step uses a different model:
DM arrives
│
├─ Skill: triage
│ llm: openai/gpt-4o-mini (fast/cheap)
│ trigger: dm
│ → Classifies the request
│
├─ chain fires ──→ Skill: deep-analysis
│ llm: anthropic/claude-sonnet-4-20250514 (powerful)
│ trigger: chain, filter: triage
│ → Performs thorough analysis
│
└─ chain fires ──→ Skill: summarize
llm: openai/gpt-4o-mini (cheap)
trigger: chain, filter: deep-analysis
→ Summarizes and DMs admin
Each skill gets its own model, temperature, and max_tokens applied independently.
Execution Parameter Resolution Per Step
For each triggered skill execution:
- Start with agent/app defaults
- Apply the skill's
llmtag (parsed asprovider/modelif slash present) - Apply the skill's
temperaturetag if present - Apply the skill's
max_tokenstag if present - Execute with those effective settings
- Restore defaults after the run
Concrete Example
Skill 1: triage (cheap fast model)
{
"kind": 31123,
"content": "## Triage\n\nClassify the incoming message:\n- If it needs deep research, use the memory_save tool to store the classification and key details.\n- If trivial, respond directly.\n\n{{message}}",
"tags": [
["d", "triage"],
["description", "Fast triage of incoming messages"],
["trigger", "dm"],
["filter", "{\"from\":\"admin\"}"],
["llm", "openai/gpt-4o-mini, cheap"],
["temperature", "0"],
["max_tokens", "200"]
]
}
Skill 2: deep-analysis (powerful model, chains from triage)
{
"kind": 31123,
"content": "## Deep Analysis\n\n{{identity}}\n\nRecall the triage classification from memory. Perform thorough analysis using available tools. Save your findings to memory for the next step.\n\nOriginal request context:\n{{message}}",
"tags": [
["d", "deep-analysis"],
["description", "Thorough analysis with powerful model"],
["trigger", "chain"],
["filter", "triage"],
["llm", "anthropic/claude-sonnet-4-20250514, best"],
["max_tokens", "2000"],
["requires_tool", "memory_recall"],
["requires_tool", "memory_save"],
["requires_tool", "nostr_query"],
["requires_skill", "identity"]
]
}
Skill 3: summarize (cheap model, chains from deep-analysis)
{
"kind": 31123,
"content": "## Summarize\n\nRecall the analysis findings from memory. Write a concise summary and DM it to admin.",
"tags": [
["d", "summarize"],
["description", "Summarize analysis and notify admin"],
["trigger", "chain"],
["filter", "deep-analysis"],
["llm", "openai/gpt-4o-mini, cheap"],
["max_tokens", "500"],
["requires_tool", "memory_recall"],
["requires_tool", "nostr_dm_send"]
]
}
Current Capabilities
| Capability | Status | Notes |
|---|---|---|
| Different model per skill | ✅ Works | Via llm tag on each skill |
| Different provider per skill | ✅ Works | provider/model format in llm tag |
| Sequential multi-step pipelines | ✅ Works | Via chain trigger type |
| Per-step temperature | ✅ Works | Via temperature tag |
| Per-step max_tokens | ✅ Works | Via max_tokens tag |
| Fallback chains per skill | ✅ Works | provider/model, provider/model, cheap |
| LLM config restore after each step | ✅ Works | Runtime saves/restores global config |
Current Limitations
1. No Direct Data Passing Between Chain Steps
Problem: Chain triggers fire with the original triggering event, not the output of the previous skill. Skill B doesn't automatically receive Skill A's output.
Current workaround: Use memory_save at the end of each step and memory_recall at the start of the next. This works but is fragile — memory is a shared scratchpad, not a structured pipeline bus.
Potential improvement: Extend the chain trigger event to include the previous skill's final LLM response text. In trigger_manager_fire_chains(), the chain event could carry a "previous_output" field that the next skill accesses via {{triggering_event}}.
2. No Conditional Branching
Problem: All chain skills matching a source d-tag fire unconditionally. You can't say "if triage classifies as X, run skill A; if Y, run skill B."
Current workaround: The chained skill can check the triggering event or memory and decide to do nothing if the condition doesn't match. But it still fires and consumes an LLM call.
Potential improvement: Add an optional chain_condition tag that the runtime evaluates before firing. Could be a simple JSON match against the previous output, or a keyword presence check.
3. No Parallel Fan-Out
Problem: Multiple chain skills matching the same source fire sequentially, not in parallel.
Current workaround: This is fine for most use cases. True parallelism would require thread-safe LLM config management.
4. Chain Depth Limit of 5
Problem: s_chain_depth in trigger_manager_fire_chains() caps at 5 levels to prevent runaway chains.
Current workaround: 5 steps is usually sufficient. For longer pipelines, the last step could use a tool to trigger a new chain externally.
Potential improvement: Make the depth limit configurable via genesis config.
5. Provider Credentials Are Global
Problem: The runtime has one set of API keys per provider. If Skill A uses anthropic/claude-sonnet-4-20250514 and Skill B uses openai/gpt-4o, both providers must be configured in the agent's LLM config. There's no per-skill credential storage.
Current workaround: Configure all needed providers in the agent's genesis config or via model_set tool. The runtime already supports provider switching via the provider field in llm_config_t.
Potential improvement: None needed for most cases — agents typically have a small number of providers configured globally.
Architecture: How It Works in Code
Trigger Execution Flow
trigger_manager fires skill
│
├─ Save current llm_config (old_cfg)
│
├─ apply_trigger_runtime_to_llm_config(trigger, &next_cfg)
│ ├─ Parse llm tag: "anthropic/claude-sonnet-4-20250514"
│ │ ├─ Set cfg->provider = "anthropic"
│ │ └─ Set cfg->model = "claude-sonnet-4-20250514"
│ ├─ Apply max_tokens if present
│ └─ Apply temperature if present
│
├─ llm_set_config(&next_cfg)
│
├─ Execute skill (agent_on_trigger)
│ ├─ Build context from triggered skills
│ ├─ Call llm_chat_with_tools_messages()
│ └─ Tool loop until completion
│
├─ Restore llm_set_config(&old_cfg)
│
└─ trigger_manager_fire_chains(source_d_tag)
├─ Find chain skills where filter == source_d_tag
├─ For each matching chain skill:
│ ├─ Save config again
│ ├─ Apply chain skill's llm override
│ ├─ Execute chain skill
│ ├─ Restore config
│ └─ Recursively fire chains (depth < 5)
└─ Done
Key Source Files
| File | Role |
|---|---|
src/trigger_manager.c |
Trigger matching, chain firing, LLM config override/restore |
src/agent.c |
agent_on_trigger() — builds context and runs LLM loop |
src/llm.c |
llm_chat_with_tools_messages() — actual LLM API call |
docs/SKILLS.md |
Skill spec including llm tag format and chain triggers |
Future Enhancements (Not Yet Implemented)
Priority 1: Chain Output Forwarding
Pass the previous skill's output to the next chain step via the triggering event:
// In trigger_manager_fire_chains():
cJSON_AddStringToObject(event, "previous_output", last_response_text);
The chained skill would access this via {{triggering_event}} in its template, seeing:
{
"type": "chain",
"source_d_tag": "triage",
"previous_output": "Classification: needs deep analysis. Key topics: ..."
}
Priority 2: Conditional Chain Firing
Add an optional chain_condition tag:
["chain_condition", "{\"previous_output_contains\":\"needs deep analysis\"}"]
The runtime would check this before firing the chain skill.
Priority 3: Configurable Chain Depth
// In genesis.jsonc:
"trigger_chain_max_depth": 10
Summary
Multi-model skill pipelines work today using per-skill llm tags and chain triggers. The main gap is data flow between steps (currently requires memory_save/recall workaround). The system is designed for this use case — each skill execution gets its own model config applied and restored — it just needs better inter-step communication to be truly seamless.