doc update

This commit is contained in:
johnmalek312
2025-10-18 02:53:57 +11:00
parent 4f2773d168
commit bca521fce2
38 changed files with 3145 additions and 10579 deletions
-4
View File
@@ -1,4 +0,0 @@
md5 0b688f460703d59bd84fe71387e626d5 v3/sdk/droid-agent.mdx
md5 47f362d52ba26155d647efec628294cf v3/sdk/base-tools.mdx
md5 2e83b80e94101d983ed52ac5ae91314f v3/sdk/adb-tools.mdx
md5 7d779482901cc5eb62288ac9c8d85199 v3/sdk/ios-tools.mdx
+15 -20
View File
@@ -23,21 +23,12 @@
"v4/quickstart"
]
},
{
"group": "Concepts",
"pages": [
"v4/concepts/architecture",
"v4/concepts/workflow-architecture",
"v4/concepts/event-streaming"
]
},
{
"group": "Guides",
"pages": [
"v4/guides/overview",
"v4/guides/cli",
"v4/guides/configuration",
"v4/guides/device-setup",
"v4/guides/cli",
"v4/guides/custom-tools-credentials",
"v4/guides/custom-variables",
"v4/guides/app-cards",
@@ -45,13 +36,26 @@
"v4/guides/telemetry-tracing"
]
},
{
"group": "Concepts",
"pages": [
"v4/concepts/overview",
"v4/concepts/agent-architecture",
"v4/concepts/scripter-agent",
"v4/concepts/shared-state",
"v4/concepts/events-and-workflows",
"v4/concepts/prompts"
]
},
{
"group": "SDK Reference",
"pages": [
"v4/sdk",
"v4/sdk/droid-agent",
"v4/sdk/adb-tools",
"v4/sdk/ios-tools",
"v4/sdk/base-tools"
"v4/sdk/base-tools",
"v4/sdk/configuration"
]
}
]
@@ -85,15 +89,6 @@
"v3/concepts/android-tools",
"v3/concepts/portal-app"
]
},
{
"group": "SDK Reference",
"pages": [
"v3/sdk/droid-agent",
"v3/sdk/adb-tools",
"v3/sdk/ios-tools",
"v3/sdk/base-tools"
]
}
]
},
+1 -1
View File
@@ -48,7 +48,7 @@ The DroidRun Portal App:
## 🚀 Installation
The DroidRun Portal App is available from the [DroidRun Portal repository](https://github.com/droidrun/droidrun-portal). For installation instructions, see the [Quickstart](/quickstart) guide.
The DroidRun Portal App is available from the [DroidRun Portal repository](https://github.com/droidrun/droidrun-portal). For installation instructions, see the [Quickstart](/v1/quickstart) guide.
## 🔧 Troubleshooting
+2 -2
View File
@@ -289,5 +289,5 @@ pip show droidrun
Now that you've got DroidRun running, you can:
- Learn about the [ReAct agent system](/concepts/agent)
- Discover all [Android interactions](/concepts/android-control)
- Learn about the [ReAct agent system](/v1/concepts/agent)
- Discover all [Android interactions](/v1/concepts/android-control)
+1 -1
View File
@@ -48,7 +48,7 @@ The DroidRun Portal App:
## 🚀 Installation
The DroidRun Portal App is available from the [DroidRun Portal repository](https://github.com/droidrun/droidrun-portal). For installation instructions, see the [Quickstart](/quickstart) guide.
The DroidRun Portal App is available from the [DroidRun Portal repository](https://github.com/droidrun/droidrun-portal). For installation instructions, see the [Quickstart](/v2/quickstart) guide.
## 🔧 Troubleshooting
+1 -1
View File
@@ -100,4 +100,4 @@ format, image_data = await tools.take_screenshot()
| `complete(success, reason)` | Finish task | None |
## Dive Deeper
You can find the SDK Reference for AdbTools [here](../sdk/adb-tools)
SDK reference documentation is available in the v4 documentation.
-425
View File
@@ -1,425 +0,0 @@
---
title: AdbTools
---
UI Actions - Core UI interaction tools for Android device control.
<a id="droidrun.tools.adb.AdbTools"></a>
## AdbTools
```python
class AdbTools(Tools)
```
Core UI interaction tools for Android device control.
<a id="droidrun.tools.adb.AdbTools.__init__"></a>
#### AdbTools.\_\_init\_\_
```python
def __init__(
serial: str | None = None,
use_tcp: bool = False,
tcp_port: int = 8080
) -> None
```
Initialize the AdbTools instance.
**Arguments**:
- `serial` - Device serial number
- `use_tcp` - Whether to use TCP communication (default: False)
- `tcp_port` - TCP port for communication (default: 8080)
<a id="droidrun.tools.adb.AdbTools.setup_tcp_forward"></a>
#### AdbTools.setup\_tcp\_forward
```python
def setup_tcp_forward() -> bool
```
Set up ADB TCP port forwarding for communication with the portal app.
**Returns**:
- `bool` - True if forwarding was set up successfully, False otherwise
<a id="droidrun.tools.adb.AdbTools.teardown_tcp_forward"></a>
#### AdbTools.teardown\_tcp\_forward
```python
def teardown_tcp_forward() -> bool
```
Remove ADB TCP port forwarding.
**Returns**:
- `bool` - True if forwarding was removed successfully, False otherwise
<a id="droidrun.tools.adb.AdbTools.__del__"></a>
#### AdbTools.\_\_del\_\_
```python
def __del__()
```
Cleanup when the object is destroyed.
<a id="droidrun.tools.adb.AdbTools.tap_by_index"></a>
#### AdbTools.tap\_by\_index
```python
def tap_by_index(index: int) -> str
```
Tap on a UI element by its index.
This function uses the cached clickable elements
to find the element with the given index and tap on its center coordinates.
**Arguments**:
- `index` - Index of the element to tap
**Returns**:
Result message
<a id="droidrun.tools.adb.AdbTools.tap_by_coordinates"></a>
#### AdbTools.tap\_by\_coordinates
```python
def tap_by_coordinates(x: int, y: int) -> bool
```
Tap on the device screen at specific coordinates.
**Arguments**:
- `x` - X coordinate
- `y` - Y coordinate
**Returns**:
Bool indicating success or failure
<a id="droidrun.tools.adb.AdbTools.tap"></a>
#### AdbTools.tap
```python
def tap(index: int) -> str
```
Tap on a UI element by its index.
This function uses the cached clickable elements from the last get_clickables call
to find the element with the given index and tap on its center coordinates.
**Arguments**:
- `index` - Index of the element to tap
**Returns**:
Result message
<a id="droidrun.tools.adb.AdbTools.swipe"></a>
#### AdbTools.swipe
```python
def swipe(
start_x: int,
start_y: int,
end_x: int,
end_y: int,
duration_ms: float = 300
) -> bool
```
Performs a straight-line swipe gesture on the device screen.
To perform a hold (long press), set the start and end coordinates to the same values and increase the duration as needed.
**Arguments**:
- `start_x` - Starting X coordinate
- `start_y` - Starting Y coordinate
- `end_x` - Ending X coordinate
- `end_y` - Ending Y coordinate
- `duration` - Duration of swipe in seconds
**Returns**:
Bool indicating success or failure
<a id="droidrun.tools.adb.AdbTools.drag"></a>
#### AdbTools.drag
```python
def drag(
start_x: int,
start_y: int,
end_x: int,
end_y: int,
duration: float = 3
) -> bool
```
Performs a straight-line drag and drop gesture on the device screen.
**Arguments**:
- `start_x` - Starting X coordinate
- `start_y` - Starting Y coordinate
- `end_x` - Ending X coordinate
- `end_y` - Ending Y coordinate
- `duration` - Duration of swipe in seconds
**Returns**:
Bool indicating success or failure
<a id="droidrun.tools.adb.AdbTools.input_text"></a>
#### AdbTools.input\_text
```python
def input_text(text: str) -> str
```
Input text on the device.
Always make sure that the Focused Element is not None before inputting text.
**Arguments**:
- `text` - Text to input. Can contain spaces, newlines, and special characters including non-ASCII.
**Returns**:
Result message
<a id="droidrun.tools.adb.AdbTools.back"></a>
#### AdbTools.back
```python
def back() -> str
```
Go back on the current view.
This presses the Android back button.
<a id="droidrun.tools.adb.AdbTools.press_key"></a>
#### AdbTools.press\_key
```python
def press_key(keycode: int) -> str
```
Press a key on the Android device.
Common keycodes:
- 3: HOME
- 4: BACK
- 66: ENTER
- 67: DELETE
**Arguments**:
- `keycode` - Android keycode to press
<a id="droidrun.tools.adb.AdbTools.start_app"></a>
#### AdbTools.start\_app
```python
def start_app(package: str, activity: str | None = None) -> str
```
Start an app on the device.
**Arguments**:
- `package` - Package name (e.g., "com.android.settings")
- `activity` - Optional activity name
<a id="droidrun.tools.adb.AdbTools.install_app"></a>
#### AdbTools.install\_app
```python
def install_app(
apk_path: str,
reinstall: bool = False,
grant_permissions: bool = True
) -> str
```
Install an app on the device.
**Arguments**:
- `apk_path` - Path to the APK file
- `reinstall` - Whether to reinstall if app exists
- `grant_permissions` - Whether to grant all permissions
<a id="droidrun.tools.adb.AdbTools.take_screenshot"></a>
#### AdbTools.take\_screenshot
```python
def take_screenshot() -> Tuple[str, bytes]
```
Take a screenshot of the device.
This function captures the current screen and adds the screenshot to context in the next message.
Also stores the screenshot in the screenshots list with timestamp for later GIF creation.
<a id="droidrun.tools.adb.AdbTools.list_packages"></a>
#### AdbTools.list\_packages
```python
def list_packages(include_system_apps: bool = False) -> List[str]
```
List installed packages on the device.
**Arguments**:
- `include_system_apps` - Whether to include system apps (default: False)
**Returns**:
List of package names
<a id="droidrun.tools.adb.AdbTools.complete"></a>
#### AdbTools.complete
```python
def complete(success: bool, reason: str = "")
```
Mark the task as finished.
**Arguments**:
- `success` - Indicates if the task was successful.
- `reason` - Reason for failure/success
<a id="droidrun.tools.adb.AdbTools.remember"></a>
#### AdbTools.remember
```python
def remember(information: str) -> str
```
Store important information to remember for future context.
This information will be extracted and included into your next steps to maintain context
across interactions. Use this for critical facts, observations, or user preferences
that should influence future decisions.
**Arguments**:
- `information` - The information to remember
**Returns**:
Confirmation message
<a id="droidrun.tools.adb.AdbTools.get_memory"></a>
#### AdbTools.get\_memory
```python
def get_memory() -> List[str]
```
Retrieve all stored memory items.
**Returns**:
List of stored memory items
<a id="droidrun.tools.adb.AdbTools.get_state"></a>
#### AdbTools.get\_state
```python
def get_state(serial: Optional[str] = None) -> Dict[str, Any]
```
Get both the a11y tree and phone state in a single call using the combined /state endpoint.
**Arguments**:
- `serial` - Optional device serial number
**Returns**:
Dictionary containing both 'a11y_tree' and 'phone_state' data
<a id="droidrun.tools.adb.AdbTools.get_a11y_tree"></a>
#### AdbTools.get\_a11y\_tree
```python
def get_a11y_tree() -> Dict[str, Any]
```
Get just the accessibility tree using the /a11y_tree endpoint.
**Returns**:
Dictionary containing accessibility tree data
<a id="droidrun.tools.adb.AdbTools.get_phone_state"></a>
#### AdbTools.get\_phone\_state
```python
def get_phone_state() -> Dict[str, Any]
```
Get just the phone state using the /phone_state endpoint.
**Returns**:
Dictionary containing phone state data
<a id="droidrun.tools.adb.AdbTools.ping"></a>
#### AdbTools.ping
```python
def ping() -> Dict[str, Any]
```
Test the TCP connection using the /ping endpoint.
**Returns**:
Dictionary with ping result
-191
View File
@@ -1,191 +0,0 @@
---
title: Tools
---
<a id="droidrun.tools.tools.Tools"></a>
## Tools
```python
class Tools(ABC)
```
Abstract base class for all tools.
This class provides a common interface for all tools to implement.
<a id="droidrun.tools.tools.Tools.ui_action"></a>
#### Tools.ui\_action
```python
def ui_action(func)
```
"
Decorator to capture screenshots and UI states for actions that modify the UI.
<a id="droidrun.tools.tools.Tools.get_state"></a>
#### Tools.get\_state
```python
def get_state() -> Dict[str, Any]
```
Get the current state of the tool.
<a id="droidrun.tools.tools.Tools.tap_by_index"></a>
#### Tools.tap\_by\_index
```python
def tap_by_index(index: int) -> str
```
Tap the element at the given index.
<a id="droidrun.tools.tools.Tools.swipe"></a>
#### Tools.swipe
```python
def swipe(
start_x: int,
start_y: int,
end_x: int,
end_y: int,
duration_ms: int = 300
) -> bool
```
Swipe from the given start coordinates to the given end coordinates.
<a id="droidrun.tools.tools.Tools.drag"></a>
#### Tools.drag
```python
def drag(
start_x: int,
start_y: int,
end_x: int,
end_y: int,
duration_ms: int = 3000
) -> bool
```
Drag from the given start coordinates to the given end coordinates.
<a id="droidrun.tools.tools.Tools.input_text"></a>
#### Tools.input\_text
```python
def input_text(text: str) -> str
```
Input the given text into a focused input field.
<a id="droidrun.tools.tools.Tools.back"></a>
#### Tools.back
```python
def back() -> str
```
Press the back button.
<a id="droidrun.tools.tools.Tools.press_key"></a>
#### Tools.press\_key
```python
def press_key(keycode: int) -> str
```
Enter the given keycode.
<a id="droidrun.tools.tools.Tools.start_app"></a>
#### Tools.start\_app
```python
def start_app(package: str, activity: str = "") -> str
```
Start the given app.
<a id="droidrun.tools.tools.Tools.take_screenshot"></a>
#### Tools.take\_screenshot
```python
def take_screenshot() -> Tuple[str, bytes]
```
Take a screenshot of the device.
<a id="droidrun.tools.tools.Tools.list_packages"></a>
#### Tools.list\_packages
```python
def list_packages(include_system_apps: bool = False) -> List[str]
```
List all packages on the device.
<a id="droidrun.tools.tools.Tools.remember"></a>
#### Tools.remember
```python
def remember(information: str) -> str
```
Remember the given information. This is used to store information in the tool's memory.
<a id="droidrun.tools.tools.Tools.get_memory"></a>
#### Tools.get\_memory
```python
def get_memory() -> List[str]
```
Get the memory of the tool.
<a id="droidrun.tools.tools.Tools.complete"></a>
#### Tools.complete
```python
def complete(success: bool, reason: str = "") -> None
```
Complete the tool. This is used to indicate that the tool has completed its task.
<a id="droidrun.tools.tools.describe_tools"></a>
#### describe\_tools
```python
def describe_tools(
tools: Tools,
exclude_tools: Optional[List[str]] = None
) -> Dict[str, Callable[..., Any]]
```
Describe the tools available for the given Tools instance.
**Arguments**:
- `tools` - The Tools instance to describe.
- `exclude_tools` - List of tool names to exclude from the description.
**Returns**:
A dictionary mapping tool names to their descriptions.
-71
View File
@@ -1,71 +0,0 @@
---
title: DroidAgent
---
DroidAgent - A wrapper class that coordinates the planning and execution of tasks
to achieve a user's goal on an Android device.
<a id="droidrun.agent.droid.droid_agent.DroidAgent"></a>
## DroidAgent
```python
class DroidAgent(Workflow)
```
A wrapper class that coordinates between PlannerAgent (creates plans) and
CodeActAgent (executes tasks) to achieve a user's goal.
<a id="droidrun.agent.droid.droid_agent.DroidAgent.__init__"></a>
#### DroidAgent.\_\_init\_\_
```python
def __init__(
goal: str,
llm: LLM,
tools: Tools,
personas: List[AgentPersona] = [DEFAULT],
max_steps: int = 15,
timeout: int = 1000,
vision: bool = False,
reasoning: bool = False,
reflection: bool = False,
enable_tracing: bool = False,
debug: bool = False,
save_trajectories: str = "none",
excluded_tools: List[str] = None,
*args,
**kwargs
)
```
Initialize the DroidAgent wrapper.
**Arguments**:
- `goal` - The user's goal or command to execute
- `llm` - The language model to use for both agents
- `max_steps` - Maximum number of steps for both agents
- `timeout` - Timeout for agent execution in seconds
- `reasoning` - Whether to use the PlannerAgent for complex reasoning (True)
or send tasks directly to CodeActAgent (False)
- `reflection` - Whether to reflect on steps the CodeActAgent did to give the PlannerAgent advice
- `enable_tracing` - Whether to enable Arize Phoenix tracing
- `debug` - Whether to enable verbose debug logging
- `save_trajectories` - Trajectory saving level. Can be:
- "none" (no saving)
- "step" (save per step)
- "action" (save per action)
- `**kwargs` - Additional keyword arguments to pass to the agents
<a id="droidrun.agent.droid.droid_agent.DroidAgent.run"></a>
#### DroidAgent.run
```python
def run(*args, **kwargs) -> WorkflowHandler
```
Run the DroidAgent workflow.
-279
View File
@@ -1,279 +0,0 @@
---
title: IOSTools
---
UI Actions - Core UI interaction tools for iOS device control.
<a id="droidrun.tools.ios.IOSTools"></a>
## IOSTools
```python
class IOSTools(Tools)
```
Core UI interaction tools for iOS device control.
<a id="droidrun.tools.ios.IOSTools.__init__"></a>
#### IOSTools.\_\_init\_\_
```python
def __init__(url: str, bundle_identifiers: List[str] = []) -> None
```
Initialize the IOSTools instance.
**Arguments**:
- `url` - iOS device URL. This is the URL of the iOS device. It is used to send requests to the iOS device.
- `bundle_identifiers` - List of bundle identifiers to include in the list of packages
<a id="droidrun.tools.ios.IOSTools.get_state"></a>
#### IOSTools.get\_state
```python
def get_state() -> List[Dict[str, Any]]
```
Get all clickable UI elements from the iOS device using accessibility API.
**Returns**:
List of dictionaries containing UI elements extracted from the device screen
<a id="droidrun.tools.ios.IOSTools.tap_by_index"></a>
#### IOSTools.tap\_by\_index
```python
def tap_by_index(index: int) -> str
```
Tap on a UI element by its index.
This function uses the cached clickable elements
to find the element with the given index and tap on its center coordinates.
**Arguments**:
- `index` - Index of the element to tap
**Returns**:
Result message
<a id="droidrun.tools.ios.IOSTools.tap"></a>
#### IOSTools.tap
```python
def tap(index: int) -> str
```
Tap on a UI element by its index.
This function uses the cached clickable elements from the last get_clickables call
to find the element with the given index and tap on its center coordinates.
**Arguments**:
- `index` - Index of the element to tap
**Returns**:
Result message
<a id="droidrun.tools.ios.IOSTools.swipe"></a>
#### IOSTools.swipe
```python
def swipe(
start_x: int,
start_y: int,
end_x: int,
end_y: int,
duration_ms: int = 300
) -> bool
```
Performs a straight-line swipe gesture on the device screen.
To perform a hold (long press), set the start and end coordinates to the same values and increase the duration as needed.
**Arguments**:
- `start_x` - Starting X coordinate
- `start_y` - Starting Y coordinate
- `end_x` - Ending X coordinate
- `end_y` - Ending Y coordinate
- `duration_ms` - Duration of swipe in milliseconds (not used in iOS API)
**Returns**:
Bool indicating success or failure
<a id="droidrun.tools.ios.IOSTools.drag"></a>
#### IOSTools.drag
```python
def drag(
start_x: int,
start_y: int,
end_x: int,
end_y: int,
duration_ms: int = 3000
) -> bool
```
Drag from the given start coordinates to the given end coordinates.
**Arguments**:
- `start_x` - Starting X coordinate
- `start_y` - Starting Y coordinate
- `end_x` - Ending X coordinate
- `end_y` - Ending Y coordinate
- `duration_ms` - Duration of swipe in milliseconds
**Returns**:
Bool indicating success or failure
<a id="droidrun.tools.ios.IOSTools.input_text"></a>
#### IOSTools.input\_text
```python
def input_text(text: str) -> str
```
Input text on the iOS device.
**Arguments**:
- `text` - Text to input. Can contain spaces, newlines, and special characters including non-ASCII.
**Returns**:
Result message
<a id="droidrun.tools.ios.IOSTools.back"></a>
#### IOSTools.back
```python
def back() -> str
```
<a id="droidrun.tools.ios.IOSTools.press_key"></a>
#### IOSTools.press\_key
```python
def press_key(keycode: int) -> str
```
Press a key on the iOS device.
iOS Key codes:
- 0: HOME
- 4: ACTION
- 5: CAMERA
**Arguments**:
- `keycode` - iOS keycode to press
<a id="droidrun.tools.ios.IOSTools.start_app"></a>
#### IOSTools.start\_app
```python
def start_app(package: str, activity: str = "") -> str
```
Start an app on the iOS device.
**Arguments**:
- `package` - Bundle identifier (e.g., "com.apple.MobileSMS")
- `activity` - Optional activity name (not used on iOS)
<a id="droidrun.tools.ios.IOSTools.take_screenshot"></a>
#### IOSTools.take\_screenshot
```python
def take_screenshot() -> Tuple[str, bytes]
```
Take a screenshot of the iOS device.
This function captures the current screen and adds the screenshot to context in the next message.
Also stores the screenshot in the screenshots list with timestamp for later GIF creation.
<a id="droidrun.tools.ios.IOSTools.list_packages"></a>
#### IOSTools.list\_packages
```python
def list_packages(include_system_apps: bool = True) -> List[str]
```
<a id="droidrun.tools.ios.IOSTools.remember"></a>
#### IOSTools.remember
```python
def remember(information: str) -> str
```
Store important information to remember for future context.
This information will be included in future LLM prompts to help maintain context
across interactions. Use this for critical facts, observations, or user preferences
that should influence future decisions.
**Arguments**:
- `information` - The information to remember
**Returns**:
Confirmation message
<a id="droidrun.tools.ios.IOSTools.get_memory"></a>
#### IOSTools.get\_memory
```python
def get_memory() -> List[str]
```
Retrieve all stored memory items.
**Returns**:
List of stored memory items
<a id="droidrun.tools.ios.IOSTools.complete"></a>
#### IOSTools.complete
```python
def complete(success: bool, reason: str = "")
```
Mark the task as finished.
**Arguments**:
- `success` - Indicates if the task was successful.
- `reason` - Reason for failure/success
+241
View File
@@ -0,0 +1,241 @@
---
title: 'Multi-Agent Architecture'
description: 'Droidrun v4 hierarchical agent system with specialized roles for planning, execution, and computation.'
---
## What is Multi-Agent Architecture?
Droidrun v4 uses a **hierarchical multi-agent system** where specialized agents work together:
- **DroidAgent**: Main orchestrator coordinating all agents
- **ManagerAgent**: Strategic planner creating task plans
- **ExecutorAgent**: Tactical actor executing atomic actions
- **CodeActAgent**: Direct code generator for simple tasks
- **ScripterAgent**: Off-device Python executor for API calls, file operations, and computations
**Location**: `droidrun/agent/droid/droid_agent.py`
## How It Works
```
DroidAgent (orchestrator)
├── Reasoning Mode: ManagerAgent → ExecutorAgent → ScripterAgent
└── Direct Mode: CodeActAgent
```
All agents share `DroidAgentState` for coordination and communicate through events.
## DroidAgent (Orchestrator)
Entry point for all tasks. Routes to appropriate agents based on mode.
```python
from droidrun.agent.droid import DroidAgent
from droidrun.config_manager import DroidrunConfig
config = DroidrunConfig()
# Reasoning mode (complex tasks)
agent = DroidAgent(reasoning=True, config=config)
# Direct mode (simple tasks)
agent = DroidAgent(reasoning=False, config=config)
result = agent.run("Send message to John")
```
## ManagerAgent (Planner)
Creates strategic plans and breaks tasks into subgoals.
**Location**: `droidrun/agent/manager/manager_agent.py:46`
```python
class ManagerPlan(BaseModel):
current_subgoal: str # Next subgoal for Executor
reasoning: str # Why this subgoal
should_finalize: bool # Task complete?
script_block: str | None # Python for ScripterAgent
full_plan: List[str] # Complete plan
```
**Configuration:**
```yaml
agent:
manager:
max_steps: 10
vision: true
llm_profiles:
manager:
provider: Anthropic
model: claude-sonnet-4
temperature: 0.7
```
## ExecutorAgent (Actor)
Executes atomic actions for each subgoal.
**Location**: `droidrun/agent/executor/executor_agent.py`
```python
class ExecutorAction(BaseModel):
action: str # "click", "type", "swipe", etc.
parameters: dict # Action parameters
reasoning: str # Why this action
class ExecutorResult(BaseModel):
success: bool # Action succeeded?
outcome: str # What happened
error_message: str | None
```
**Configuration:**
```yaml
agent:
executor:
max_steps: 5
vision: true
llm_profiles:
executor:
provider: OpenAI
model: gpt-4o
temperature: 0.3
```
## CodeActAgent (Direct Executor)
Generates Python code using atomic actions (no planning overhead).
**Location**: `droidrun/agent/codeact/codeact_agent.py`
```python
# Available functions in CodeAct
click(index: int)
long_press(index: int)
type(text: str, index: int = None)
swipe(coordinate: tuple, coordinate2: tuple)
system_button(button: str)
open_app(text: str)
get_state() -> dict
take_screenshot() -> str
remember(information: str)
complete(success: bool, reason: str)
```
**Configuration:**
```yaml
agent:
codeact:
max_steps: 15
vision: false
safe_execution:
enabled: true
llm_profiles:
codeact:
provider: GoogleGenAI
model: models/gemini-2.0-flash-exp
```
## ScripterAgent (Python Executor)
Executes off-device Python for API calls, file operations, data processing, and computations.
**Location**: `droidrun/agent/scripter/`
Triggered when Manager delegates tasks requiring off-device computation. ScripterAgent is a **ReAct agent** that iteratively generates and executes Python code, then returns a final message to Manager.
```python
# Manager delegates with context + task
"""
User needs weather in San Francisco for clothing decision.
Task: Fetch current weather and report temperature + conditions
API: https://api.weather.com/forecast?city=San Francisco
"""
# ScripterAgent (ReAct loop):
# 1. Generates code
import requests
response = requests.get("https://api.weather.com/forecast",
params={"city": "San Francisco"})
print(response.json())
# 2. Observes output: {'temp': 62, 'description': 'Partly cloudy'}
# 3. Returns message to Manager:
"The weather in San Francisco is 62°F with partly cloudy conditions."
```
**Configuration:**
```yaml
agent:
scripter:
max_steps: 10
safe_execution:
enabled: true
allowed_modules:
- datetime
- json
- requests
```
## Agent Coordination
### Shared State
All agents read/write `DroidAgentState`:
```python
state = DroidAgentState(
task="Book flight",
action_history=[],
visited_packages=[],
error_count=0,
scripter_results={},
manager_plan="",
executor_feedback="",
step_count=0
)
```
### Event Flow (Reasoning Mode)
```
StartEvent
↓
ManagerInputEvent → run_manager()
↓
ManagerPlanEvent → handle_manager_plan()
↓
ExecutorInputEvent → run_executor()
↓
ExecutorResultEvent → handle_executor_result()
↓
[loop or finalize]
↓
FinalizeEvent → finalize()
↓
ResultEvent (StopEvent)
```
## Quick Reference
| Agent | Role | Best For | Config Key |
|-------|------|----------|------------|
| DroidAgent | Orchestrator | Entry point | `agent.*` |
| ManagerAgent | Planner | Strategy, recovery | `agent.manager.*` |
| ExecutorAgent | Actor | Action execution | `agent.executor.*` |
| CodeActAgent | Direct | Simple tasks | `agent.codeact.*` |
| ScripterAgent | Python Executor | APIs, files, data | `agent.scripter.*` |
## Related Topics
- [Reasoning Mode](./reasoning-mode) - Manager → Executor workflow
- [Direct Mode](./direct-mode) - CodeActAgent workflow
- [ScripterAgent](./scripter-agent) - Off-device computation
- [Shared State](./shared-state) - DroidAgentState coordination
- [Configuration](./configuration) - Per-agent LLM profiles
File diff suppressed because it is too large Load Diff
-975
View File
@@ -1,975 +0,0 @@
---
title: "Event Streaming"
description: "Understanding DroidRun event system for real-time monitoring and custom handlers"
---
## Overview
DroidRun features a comprehensive event streaming architecture built on [LlamaIndex Workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/). Events flow through the system in real-time, enabling you to monitor agent execution, build custom UIs, and integrate with external systems.
**Quick Start:**
```python
from droidrun import DroidAgent, ResultEvent
from droidrun.config_manager.config_manager import DroidRunConfig
config = DroidRunConfig()
agent = DroidAgent(goal="Your task", config=config)
handler = agent.run()
# Stream events in real-time
async for event in handler.stream_events():
print(f"Event: {event.__class__.__name__}")
# Get final result (ResultEvent with success, reason, steps, structured_output)
result: ResultEvent = await handler
print(f"Success: {result.success}, Reason: {result.reason}")
```
### Two-Tier Event System
DroidRun uses a dual-layer event architecture:
1. **Coordination Events** (`droidrun/agent/droid/events.py`) - Lightweight routing events for workflow orchestration between DroidAgent and child agents
2. **Internal Events** (agent-specific `events.py` files) - Rich debugging events with full metadata, streamed to external consumers
```python
# Coordination Event (workflow routing only, NOT streamed)
class ExecutorInputEvent(Event):
current_subgoal: str # Minimal data for routing
# Internal Event (streamed to frontend/logs with full context)
class ExecutorInternalActionEvent(Event):
action_json: str # Raw JSON of selected action
thought: str # LLM's reasoning process
description: str # Human-readable description
```
**Key distinction:**
- **Coordination events** trigger workflow steps (e.g., `ManagerInputEvent` → runs Manager)
- **Internal events** are streamed for monitoring (e.g., `ManagerInternalPlanEvent` → shows plan details)
- **ResultEvent** is the final return value from `agent.run()` (includes success, reason, steps, structured_output)
- **StopEvent** is a LlamaIndex internal event (filtered, never streamed to consumers)
## Event Categories
### Coordination Events
Located in `/droidrun/agent/droid/events.py`, these events control workflow flow between DroidAgent and child agents:
| Event | Purpose | Fields |
|-------|---------|--------|
| `ManagerInputEvent` | Trigger Manager planning | None (signal only) |
| `ManagerPlanEvent` | Manager → DroidAgent routing | `plan`, `current_subgoal`, `thought`, `manager_answer` |
| `ExecutorInputEvent` | Trigger Executor action | `current_subgoal` |
| `ExecutorResultEvent` | Executor → DroidAgent routing | `action`, `outcome`, `error`, `summary`, `full_response` |
| `CodeActExecuteEvent` | Trigger CodeAct execution | `task` (Task object) |
| `CodeActResultEvent` | CodeAct → DroidAgent routing | `success`, `reason`, `task` |
| `ScripterExecutorInputEvent` | Trigger Scripter workflow | `task` (str) |
| `ScripterExecutorResultEvent` | Scripter → DroidAgent routing | `task`, `message`, `success`, `code_executions` |
| `FinalizeEvent` | Signal task completion | `success`, `reason` |
| `ResultEvent` | Final workflow result | `success`, `reason`, `steps`, `structured_output` |
### Internal Events (Streamed)
These events are written to the event stream and contain full debugging metadata. They are located in agent-specific event files:
#### Manager Events
```python
# File: droidrun/agent/manager/events.py
class ManagerThinkingEvent(Event):
"""Manager is analyzing state"""
pass
class ManagerInternalPlanEvent(Event):
"""Manager created a plan (with full metadata)"""
plan: str
current_subgoal: str
thought: str
manager_answer: str = ""
memory_update: str = "" # Debugging: LLM's memory additions
```
#### Executor Events
```python
# File: droidrun/agent/executor/events.py
class ExecutorInternalActionEvent(Event):
"""Executor selected an action (with reasoning)"""
action_json: str # Raw JSON string
thought: str # LLM's reasoning process
description: str # Human-readable description
full_response: str = "" # Full LLM response for development
class ExecutorInternalResultEvent(Event):
"""Executor completed action (with full debug info)"""
action: Dict
outcome: bool
error: str
summary: str
thought: str = ""
action_json: str = ""
full_response: str = "" # Full LLM response for development
```
#### CodeAct Events
```python
# File: droidrun/agent/codeact/events.py
class TaskInputEvent(Event):
"""Task input received"""
input: list[ChatMessage]
class TaskThinkingEvent(Event):
"""CodeAct is reasoning"""
thoughts: Optional[str] = None
code: Optional[str] = None
usage: Optional[UsageResult] = None
class TaskExecutionEvent(Event):
"""Executing code"""
code: str
globals: dict[str, str] = {}
locals: dict[str, str] = {}
class TaskExecutionResultEvent(Event):
"""Code execution result"""
output: str
class TaskEndEvent(Event):
"""Task finished"""
success: bool
reason: str
class EpisodicMemoryEvent(Event):
"""Episodic memory snapshot"""
episodic_memory: EpisodicMemory
```
#### Scripter Events
```python
# File: droidrun/agent/scripter/events.py
class ScripterInputEvent(Event):
"""Scripter input"""
input: List # List of ChatMessages
class ScripterThinkingEvent(Event):
"""Scripter thinking"""
thoughts: str
code: Optional[str] = None
full_response: str = ""
class ScripterExecutionEvent(Event):
"""Scripter executing code"""
code: str
class ScripterExecutionResultEvent(Event):
"""Scripter execution result"""
output: str
class ScripterEndEvent(Event):
"""Scripter finished"""
message: str
success: bool
code_executions: int = 0
```
#### Common Events
```python
# File: droidrun/agent/common/events.py
class ScreenshotEvent(Event):
"""Screenshot captured"""
screenshot: bytes
class RecordUIStateEvent(Event):
"""UI state recorded"""
ui_state: list[Dict[str, Any]]
# Macro events for coordinate-based actions
class TapActionEvent(MacroEvent):
x: int
y: int
element_index: int = None
element_text: str = ""
element_bounds: str = ""
class SwipeActionEvent(MacroEvent):
start_x: int
start_y: int
end_x: int
end_y: int
duration_ms: int
class DragActionEvent(MacroEvent):
start_x: int
start_y: int
end_x: int
end_y: int
duration_ms: int
class InputTextActionEvent(MacroEvent):
text: str
class KeyPressActionEvent(MacroEvent):
keycode: int
key_name: str = ""
class StartAppEvent(MacroEvent):
package: str
activity: str = None
```
### Telemetry Events
Located in `/droidrun/telemetry/events.py`, these capture high-level metrics:
```python
class DroidAgentInitEvent(TelemetryEvent):
"""Agent initialization"""
goal: str
llms: Dict[str, str]
tools: str
max_steps: int
timeout: int
vision: Dict[str, bool]
reasoning: bool
enable_tracing: bool
debug: bool
save_trajectories: str = "none"
runtype: str = "developer"
class PackageVisitEvent(TelemetryEvent):
"""App/package visited"""
package_name: str
activity_name: str
step_number: int
class DroidAgentFinalizeEvent(TelemetryEvent):
"""Agent finalized"""
success: bool
reason: str
steps: int
unique_packages_count: int
unique_activities_count: int
```
## Capturing Events
### Basic Event Streaming
The simplest way to capture events is through the workflow handler's `stream_events()` method:
```python
from droidrun import DroidAgent, ResultEvent
from droidrun.config_manager.config_manager import DroidRunConfig
# Initialize agent with config
config = DroidRunConfig()
agent = DroidAgent(
goal="Open Settings and enable WiFi",
config=config
)
# Run and capture events
handler = agent.run()
async for event in handler.stream_events():
print(f"Event: {event.__class__.__name__}")
# Access event-specific fields safely
if hasattr(event, 'thought'):
print(f" Thought: {event.thought}")
if hasattr(event, 'action'):
print(f" Action: {event.action}")
# Get final result
result: ResultEvent = await handler
print(f"Success: {result.success}")
print(f"Reason: {result.reason}")
print(f"Steps: {result.steps}")
```
### Building a Custom Event Handler
Create rich monitoring experiences by building custom handlers:
```python
import logging
from typing import List
from rich.console import Console
from rich.live import Live
from rich.panel import Panel
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.agent.manager.events import ManagerThinkingEvent, ManagerInternalPlanEvent
from droidrun.agent.executor.events import ExecutorInternalActionEvent, ExecutorInternalResultEvent
from droidrun.agent.codeact.events import TaskThinkingEvent, TaskExecutionResultEvent
from droidrun.agent.common.events import ScreenshotEvent, RecordUIStateEvent
from droidrun.agent.droid.events import FinalizeEvent
class CustomEventHandler(logging.Handler):
def __init__(self):
super().__init__()
self.console = Console()
self.events: List[str] = []
def handle_event(self, event):
"""Process different event types"""
# Manager events
if isinstance(event, ManagerThinkingEvent):
self.console.print("🧠 [blue]Manager analyzing state...[/blue]")
elif isinstance(event, ManagerInternalPlanEvent):
self.console.print(f"📋 [green]Plan created:[/green] {event.current_subgoal}")
if event.memory_update:
self.console.print(f" 💾 Memory: {event.memory_update[:100]}...")
# Executor events
elif isinstance(event, ExecutorInternalActionEvent):
self.console.print(f"🎯 [yellow]Action:[/yellow] {event.description}")
self.console.print(f" 💭 Thought: {event.thought[:100]}...")
elif isinstance(event, ExecutorInternalResultEvent):
status = "✅" if event.outcome else "❌"
self.console.print(f"{status} {event.summary}")
if not event.outcome:
self.console.print(f" ⚠️ Error: {event.error}")
# CodeAct events
elif isinstance(event, TaskThinkingEvent):
if event.thoughts:
self.console.print(f"💭 [cyan]Thinking:[/cyan] {event.thoughts[:150]}...")
if event.code:
self.console.print("💻 [magenta]Executing code[/magenta]")
elif isinstance(event, TaskExecutionResultEvent):
self.console.print(f"📤 Result: {event.output[:100]}...")
# Screenshot & UI state
elif isinstance(event, ScreenshotEvent):
self.console.print("📸 Screenshot captured")
elif isinstance(event, RecordUIStateEvent):
element_count = len(event.ui_state)
self.console.print(f"📱 UI state: {element_count} elements")
# Completion
elif isinstance(event, FinalizeEvent):
status = "🎉 Success" if event.success else "❌ Failed"
self.console.print(f"{status}: {event.reason}")
# Use the custom handler
handler_instance = CustomEventHandler()
config = DroidRunConfig()
agent = DroidAgent(goal="Take a screenshot", config=config)
workflow_handler = agent.run()
async for event in workflow_handler.stream_events():
handler_instance.handle_event(event)
result: ResultEvent = await workflow_handler
print(f"\nFinal result: {result.success} - {result.reason}")
```
### Real-World Example: LogHandler
DroidRun's CLI uses a sophisticated event handler with Rich UI. Here's the simplified version from `/droidrun/cli/logs.py`:
```python
class LogHandler(logging.Handler):
def __init__(self, goal: str, rich_text: bool = True):
super().__init__()
self.goal = goal
self.current_step = "Initializing..."
self.is_completed = False
self.is_success = False
if rich_text:
self.console = Console()
self.logs: List[str] = []
def handle_event(self, event):
"""Handle streaming events from agent workflow"""
# Manager events
if isinstance(event, ManagerThinkingEvent):
self.current_step = "Manager analyzing state..."
elif isinstance(event, ManagerInternalPlanEvent):
self.current_step = "Plan created"
if event.current_subgoal:
logger.info(f"📋 Next step: {event.current_subgoal[:150]}")
# Executor events
elif isinstance(event, ExecutorInternalActionEvent):
self.current_step = "Selecting action..."
logger.info(f"🎯 Action: {event.description}")
elif isinstance(event, ExecutorInternalResultEvent):
if event.outcome:
self.current_step = "Action completed"
logger.info(f"✅ {event.summary}")
else:
self.current_step = "Action failed"
logger.info(f"❌ {event.summary} ({event.error})")
# CodeAct events
elif isinstance(event, TaskThinkingEvent):
if event.thoughts:
logger.info(f"🧠 Thinking: {event.thoughts[:150]}")
if event.code:
logger.info("💻 Executing action code")
elif isinstance(event, TaskExecutionResultEvent):
if "Error" in str(event.output):
logger.info(f"❌ Action error: {event.output[:100]}")
else:
logger.info(f"⚡ Result: {event.output[:100]}")
# Finalization
elif isinstance(event, FinalizeEvent):
self.is_completed = True
self.is_success = event.success
if event.success:
logger.info(f"🎉 Goal achieved: {event.reason}")
else:
logger.info(f"❌ Goal failed: {event.reason}")
# Usage in CLI
log_handler = LogHandler(goal=goal, rich_text=True)
logger.addHandler(log_handler)
with log_handler.render(): # Rich Live rendering
handler = agent.run()
async for event in handler.stream_events():
log_handler.handle_event(event)
result: ResultEvent = await handler
```
## Shared State Access
DroidAgent maintains shared state (`DroidAgentState`) accessible across all agents:
```python
from droidrun.agent.droid.events import DroidAgentState
# The shared state (accessed via agent.shared_state)
state = DroidAgentState(
# Task context
instruction="Your goal here",
step_number=0,
# Device state
formatted_device_state="...", # Complete UI hierarchy text
a11y_tree=[...], # Raw accessibility tree (list of dicts)
phone_state={...}, # Phone metadata (dict)
screenshot=bytes, # Current screenshot
width=1080, # Screen width
height=1920, # Screen height
focused_text="...", # Currently focused element text
# App Cards
app_card="...", # Current app-specific instructions
current_package_name="...", # Current app package
current_activity_name="...", # Current activity
# Action tracking
action_history=[...], # All actions taken (list of dicts)
action_pool=[...], # Raw action JSONs
summary_history=[...], # Action summaries (list of str)
action_outcomes=[...], # Success/failure bools (list of bool)
error_descriptions=[...], # Error messages (list of str)
last_action={...}, # Most recent action (dict)
last_summary="...", # Most recent summary
last_action_thought="...", # Most recent thought
# Planning
plan="...", # Current plan
current_subgoal="...", # Active subgoal
finish_thought="...", # Completion reasoning
progress_status="...", # Current progress
manager_answer="...", # Answer for answer-type tasks
memory="...", # Agent memory
message_history=[...], # Chat history
# Error handling
error_flag_plan=False, # Error escalation flag
err_to_manager_thresh=2, # Error threshold
# App tracking
visited_packages=set(), # Visited apps
visited_activities=set(), # Visited activities
# Script execution
scripter_history=[...], # Script results (list of dicts)
last_scripter_message="...", # Most recent scripter response
last_scripter_success=True, # Most recent scripter status
# Device state comparison
previous_formatted_device_state="...", # For before/after comparison
# Output and metadata
output_dir="", # Output directory path
user_id=None, # Optional user ID
# Custom data
custom_variables={...}, # User-defined variables
)
# Access in events
async for event in handler.stream_events():
if isinstance(event, ManagerInternalPlanEvent):
# State is shared across all agents
current_plan = agent.shared_state.plan
current_step = agent.shared_state.step_number
visited_apps = agent.shared_state.visited_packages
```
## Event Flow by Mode
### Direct Execution Mode (`reasoning=False`)
```mermaid
graph LR
Start[StartEvent] --> CodeAct[CodeActExecuteEvent]
CodeAct --> Input[TaskInputEvent]
Input --> Think[TaskThinkingEvent]
Think --> Exec[TaskExecutionEvent]
Exec --> Result[TaskExecutionResultEvent]
Result --> Input
Result --> End[TaskEndEvent]
End --> Final[FinalizeEvent]
Final --> Return[ResultEvent]
```
**Events emitted:**
1. `StartEvent` → DroidAgent starts (streamed)
2. `CodeActExecuteEvent` → CodeAct agent triggered
3. `TaskInputEvent` → Task input prepared
4. `TaskThinkingEvent` → LLM thinking (thoughts + code)
5. `TaskExecutionEvent` → Code executing
6. `TaskExecutionResultEvent` → Execution result
7. Loop back to step 3 or proceed to 8
8. `TaskEndEvent` → Task complete
9. `FinalizeEvent` → DroidAgent finalizing (streamed)
10. Optional: If `output_model` is provided, StructuredOutputAgent runs to extract structured data
11. `ResultEvent` → Final result returned (success, reason, steps, structured_output)
### Reasoning Mode (`reasoning=True`)
```mermaid
graph LR
Start[StartEvent] --> Manager[ManagerInputEvent]
Manager --> Think[ManagerThinkingEvent]
Think --> Plan[ManagerInternalPlanEvent]
Plan --> Route{Route}
Route -->|action| Exec[ExecutorInputEvent]
Route -->|script| Script[ScripterExecutorInputEvent]
Route -->|answer| Final[FinalizeEvent]
Exec --> ExecAct[ExecutorInternalActionEvent]
ExecAct --> ExecRes[ExecutorInternalResultEvent]
ExecRes --> Manager
Script --> ScriptEnd[ScripterEndEvent]
ScriptEnd --> Manager
Final --> Return[ResultEvent]
```
**Events emitted (per cycle):**
**Manager Phase:**
1. `ManagerInputEvent` → Manager triggered
2. `ManagerThinkingEvent` → Manager analyzing
3. `ManagerInternalPlanEvent` → Plan created (with thought, subgoal)
**Executor Phase (if action needed):**
4. `ExecutorInputEvent` → Executor triggered
5. `ExecutorInternalActionEvent` → Action selected (with thought)
6. `ExecutorInternalResultEvent` → Action result
7. Loop back to Manager (step 1)
**Scripter Phase (if `<script>` tag):**
4. `ScripterExecutorInputEvent` → Scripter triggered
5. `ScripterInputEvent` → Scripter input prepared
6. `ScripterThinkingEvent` → Scripter thinking
7. `ScripterExecutionEvent` → Code executing
8. `ScripterExecutionResultEvent` → Execution result
9. `ScripterEndEvent` → Scripter complete
10. Loop back to Manager (step 1)
**Completion:**
- `FinalizeEvent` → Task complete (when Manager provides answer, streamed)
- Optional: If `output_model` is provided, StructuredOutputAgent runs to extract structured data from the final answer
- `ResultEvent` → Final result returned (success, reason, steps, structured_output)
## Nested Event Handling
DroidAgent streams nested workflow events to its parent context. This is implemented in the `handle_stream_event()` method:
```python
# In DroidAgent.handle_stream_event()
def handle_stream_event(self, ev: Event, ctx: Context):
"""Route nested events to parent stream or handle internally"""
# Internal handling (don't stream upstream)
if isinstance(ev, EpisodicMemoryEvent):
self.current_episodic_memory = ev.episodic_memory
return
# Skip StopEvent (workflow-internal coordination only)
if not isinstance(ev, StopEvent):
ctx.write_event_to_stream(ev) # Stream to parent
# Trajectory tracking for different event types
if isinstance(ev, ScreenshotEvent):
self.trajectory.screenshots.append(ev.screenshot)
elif isinstance(ev, MacroEvent):
self.trajectory.macro.append(ev)
elif isinstance(ev, RecordUIStateEvent):
self.trajectory.ui_states.append(ev.ui_state)
else:
self.trajectory.events.append(ev)
```
**Key behaviors:**
- All child workflow events (Manager, Executor, CodeAct, Scripter) are automatically streamed to the parent
- `EpisodicMemoryEvent` is captured internally but not streamed (used for memory state tracking)
- `StopEvent` is filtered out (workflow-internal coordination only, not for external consumers)
- Events are captured for trajectory recording when `save_trajectory` is enabled
## Best Practices
### 1. Event Type Checking
Always use `isinstance()` for type-safe event handling:
```python
if isinstance(event, ManagerInternalPlanEvent):
# TypeScript-safe: event.plan is guaranteed to exist
print(event.plan)
```
### 2. Graceful Field Access
Use `hasattr()` for optional fields:
```python
if hasattr(event, 'memory_update') and event.memory_update:
print(f"Memory: {event.memory_update}")
```
### 3. Event Filtering
Filter events by category for focused monitoring:
```python
# Only planning events
async for event in handler.stream_events():
if isinstance(event, (ManagerThinkingEvent, ManagerInternalPlanEvent)):
# Handle planning events
pass
# Only action events
async for event in handler.stream_events():
if isinstance(event, (ExecutorInternalActionEvent, ExecutorInternalResultEvent)):
# Handle action events
pass
```
### 4. Performance Considerations
For high-frequency events, use buffering:
```python
import asyncio
from collections import deque
event_buffer = deque(maxlen=100)
async def buffer_events():
async for event in handler.stream_events():
event_buffer.append(event)
# Process in batches
if len(event_buffer) >= 10:
await process_batch(list(event_buffer))
event_buffer.clear()
```
### 5. Error Handling
Always wrap event handling in try-except:
```python
async for event in handler.stream_events():
try:
handler_instance.handle_event(event)
except Exception as e:
logger.error(f"Event handling failed: {e}", exc_info=True)
# Continue processing other events
continue
```
## Advanced Patterns
### Event Aggregation
Collect events for post-execution analysis:
```python
from collections import defaultdict
class EventAggregator:
def __init__(self):
self.events_by_type = defaultdict(list)
self.timeline = []
def capture(self, event):
event_type = event.__class__.__name__
self.events_by_type[event_type].append(event)
self.timeline.append((time.time(), event))
def get_stats(self):
return {
"total_events": len(self.timeline),
"by_type": {k: len(v) for k, v in self.events_by_type.items()},
"duration": self.timeline[-1][0] - self.timeline[0][0] if self.timeline else 0,
}
# Usage
aggregator = EventAggregator()
async for event in handler.stream_events():
aggregator.capture(event)
stats = aggregator.get_stats()
print(f"Processed {stats['total_events']} events in {stats['duration']:.2f}s")
```
### Real-Time Webhooks
Stream events to external systems:
```python
import httpx
import time
class WebhookStreamer:
def __init__(self, webhook_url: str):
self.webhook_url = webhook_url
self.client = httpx.AsyncClient()
async def send_event(self, event):
payload = {
"type": event.__class__.__name__,
"timestamp": time.time(),
"data": event.model_dump() if hasattr(event, 'model_dump') else str(event),
}
try:
await self.client.post(self.webhook_url, json=payload)
except Exception as e:
print(f"Webhook failed: {e}")
async def close(self):
await self.client.aclose()
# Usage
from droidrun.agent.manager.events import ManagerInternalPlanEvent
from droidrun.agent.executor.events import ExecutorInternalResultEvent
from droidrun.agent.droid.events import FinalizeEvent
streamer = WebhookStreamer("https://api.example.com/events")
try:
async for event in handler.stream_events():
# Filter important events
if isinstance(event, (ManagerInternalPlanEvent, ExecutorInternalResultEvent, FinalizeEvent)):
await streamer.send_event(event)
finally:
await streamer.close()
```
### Event Replay
Record and replay events for debugging:
```python
import pickle
class EventRecorder:
def __init__(self, filepath: str):
self.filepath = filepath
self.events = []
def record(self, event):
self.events.append(event)
def save(self):
with open(self.filepath, 'wb') as f:
pickle.dump(self.events, f)
@staticmethod
def load(filepath: str):
with open(filepath, 'rb') as f:
return pickle.load(f)
def replay(self, handler_func):
"""Replay recorded events through a handler"""
for event in self.events:
handler_func(event)
# Record
recorder = EventRecorder("session.pkl")
async for event in handler.stream_events():
recorder.record(event)
recorder.save()
# Replay
events = EventRecorder.load("session.pkl")
for event in events:
print(f"Replaying: {event.__class__.__name__}")
```
## Integration Examples
### FastAPI WebSocket Streaming
```python
from fastapi import FastAPI, WebSocket
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
import json
app = FastAPI()
@app.websocket("/agent/stream")
async def agent_stream(websocket: WebSocket):
await websocket.accept()
# Get goal from client
goal = await websocket.receive_text()
# Create agent with config
config = DroidRunConfig()
agent = DroidAgent(goal=goal, config=config)
handler = agent.run()
# Stream events to WebSocket
try:
async for event in handler.stream_events():
event_data = {
"type": event.__class__.__name__,
"data": event.model_dump() if hasattr(event, 'model_dump') else {},
}
await websocket.send_text(json.dumps(event_data))
# Send final result
result: ResultEvent = await handler
await websocket.send_text(json.dumps({
"type": "result",
"data": {
"success": result.success,
"reason": result.reason,
"steps": result.steps,
}
}))
except Exception as e:
await websocket.send_text(json.dumps({"type": "error", "message": str(e)}))
finally:
await websocket.close()
```
### Discord Bot Integration
```python
import discord
import time
from discord.ext import commands
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.agent.manager.events import ManagerInternalPlanEvent
from droidrun.agent.executor.events import ExecutorInternalActionEvent
bot = commands.Bot(command_prefix="!")
@bot.command()
async def automate(ctx, *, goal: str):
"""Run DroidRun automation and stream status to Discord"""
status_msg = await ctx.send(f"🤖 Starting automation: {goal}")
config = DroidRunConfig()
agent = DroidAgent(goal=goal, config=config)
handler = agent.run()
last_update = time.time()
async for event in handler.stream_events():
# Update every 2 seconds to avoid rate limits
if time.time() - last_update > 2:
if isinstance(event, ManagerInternalPlanEvent):
await status_msg.edit(content=f"📋 Planning: {event.current_subgoal[:100]}")
elif isinstance(event, ExecutorInternalActionEvent):
await status_msg.edit(content=f"⚡ Action: {event.description[:100]}")
last_update = time.time()
result: ResultEvent = await handler
if result.success:
await status_msg.edit(content=f"✅ Success: {result.reason[:1000]}")
else:
await status_msg.edit(content=f"❌ Failed: {result.reason[:1000]}")
```
## Troubleshooting
### Events Not Streaming
**Problem:** No events appearing in stream
**Solution:**
```python
# Ensure you're awaiting the async iterator
async for event in handler.stream_events(): # ✅ Correct
print(event)
# Not:
for event in handler.stream_events(): # ❌ Wrong
print(event)
```
### Missing Event Fields
**Problem:** `AttributeError: 'Event' object has no attribute 'X'`
**Solution:**
```python
# Always check field existence
if hasattr(event, 'thought') and event.thought:
print(event.thought)
# Or use getattr with default
thought = getattr(event, 'thought', 'No thought provided')
```
### Nested Events Not Appearing
**Problem:** Child workflow events not visible
**Solution:**
```python
# DroidAgent automatically streams nested events
# Just ensure you're iterating the top-level handler
agent = DroidAgent(...)
handler = agent.run() # Top-level workflow
# This captures ALL events (DroidAgent + children)
async for event in handler.stream_events():
print(event) # Will include Manager, Executor, CodeAct events
```
## See Also
- [Agent Architecture](/docs/v4/concepts/agent) - Understanding the multi-agent system
- [LlamaIndex Workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/) - Underlying workflow framework
- [Configuration](/docs/v4/guides/configuration) - Configuring agent behavior
+165
View File
@@ -0,0 +1,165 @@
---
title: 'Event Streaming'
description: 'How to consume real-time events from DroidAgent execution.'
---
## Overview
Droidrun provides **real-time event streaming** that gives you visibility into agent execution as it happens. This allows you to build UIs, logging systems, or monitoring tools that react to agent actions in real-time.
Under the hood, Droidrun uses [llama-index workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/) - an event-driven orchestration system that powers the agent architecture.
## Basic Usage
```python
from droidrun.agent.droid import DroidAgent
# Create and run agent
agent = DroidAgent(goal="Open Gmail and check inbox", config=config)
handler = agent.run()
# Stream events in real-time
async for event in handler.stream_events():
if isinstance(event, ManagerInternalPlanEvent):
print(f"📋 Plan: {event.plan}")
print(f"🎯 Current subgoal: {event.current_subgoal}")
elif isinstance(event, ExecutorInternalActionEvent):
print(f"⚡ Action: {event.description}")
print(f"💭 Thought: {event.thought}")
elif isinstance(event, ScreenshotEvent):
save_screenshot(event.screenshot, f"step_{event.step}.png")
elif isinstance(event, CodeGenerationEvent):
print(f"🐍 Generated code (step {event.step_number}):")
print(event.code)
# Wait for final result
result = await handler
print(f"✅ Success: {result.success}")
print(f"📝 Reason: {result.reason}")
```
## Event Types
### Planning Events
**ManagerInternalPlanEvent** - Emitted when Manager creates/updates a plan:
```python
class ManagerInternalPlanEvent(Event):
plan: str # Full task plan with subgoals
current_subgoal: str # Current subgoal being executed
thought: str # Manager's reasoning
manager_answer: str # Direct answer (if task is complete)
```
### Execution Events
**ExecutorInternalActionEvent** - Emitted when Executor selects an action:
```python
class ExecutorInternalActionEvent(Event):
action_json: str # JSON representation of selected action
thought: str # Executor's reasoning
description: str # Human-readable action description
```
**CodeGenerationEvent** - Emitted when CodeAct generates code:
```python
class CodeGenerationEvent(Event):
code: str # Generated Python code
step_number: int # Current step in execution
```
**CodeExecutionResultEvent** - Emitted after code execution:
```python
class CodeExecutionResultEvent(Event):
success: bool # Whether execution succeeded
output: str # Execution output or error message
```
### Visual Events
**ScreenshotEvent** - Emitted when a screenshot is captured:
```python
class ScreenshotEvent(Event):
screenshot: bytes # PNG image data
step: int # Step number
```
**AccessibilityTreeEvent** - Emitted when UI tree is captured:
```python
class AccessibilityTreeEvent(Event):
tree: str # Accessibility tree dump
step: int # Step number
```
## Common Patterns
### Building a Live UI
```python
async def run_with_ui(goal: str):
agent = DroidAgent(goal=goal, config=config)
handler = agent.run()
async for event in handler.stream_events():
if isinstance(event, ManagerInternalPlanEvent):
ui.update_plan(event.plan)
ui.update_current_step(event.current_subgoal)
elif isinstance(event, ExecutorInternalActionEvent):
ui.add_action_log(event.description, event.thought)
elif isinstance(event, ScreenshotEvent):
ui.update_screenshot(event.screenshot)
result = await handler
ui.show_completion(result.success, result.reason)
```
### Logging and Monitoring
```python
import logging
logger = logging.getLogger("droidrun.monitor")
async def monitor_execution(goal: str):
agent = DroidAgent(goal=goal, config=config)
handler = agent.run()
start_time = time.time()
action_count = 0
async for event in handler.stream_events():
if isinstance(event, ExecutorInternalActionEvent):
action_count += 1
logger.info(f"Action {action_count}: {event.description}")
elif isinstance(event, CodeExecutionResultEvent):
if not event.success:
logger.error(f"Code execution failed: {event.output}")
result = await handler
duration = time.time() - start_time
logger.info(f"Task completed in {duration:.2f}s with {action_count} actions")
logger.info(f"Result: {result.success} - {result.reason}")
```
## Notes
- Events are **streamed in real-time** as the agent executes
- Not all events are emitted in every execution (depends on mode and actions)
- **Reasoning mode** emits `ManagerInternalPlanEvent` and `ExecutorInternalActionEvent`
- **Direct mode** emits `CodeGenerationEvent` and `CodeExecutionResultEvent`
- All events are **Pydantic models** with full type safety
- The `handler` object is **async** - always use `await handler` to get the final result
## Learn More
- [LlamaIndex Workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/) - The underlying orchestration system
- [Agent Architecture](./agent-architecture) - Multi-agent system overview
- [Reasoning Mode](./reasoning-mode) - Manager/Executor workflow details
- [Direct Mode](./direct-mode) - CodeAct workflow details
+107
View File
@@ -0,0 +1,107 @@
---
title: 'Architecture Overview'
description: 'Understanding Droidrun v4 multi-agent system for device automation.'
---
## What is Droidrun?
Droidrun uses a **multi-agent architecture** where specialized agents work together to complete tasks. Instead of one agent doing everything, different agents handle planning, execution, and computation.
### Two Execution Modes
- **Reasoning Mode** (`reasoning=True`): Manager plans, Executor acts. Best for complex tasks.
- **Direct Mode** (`reasoning=False`): CodeActAgent executes immediately. Best for simple tasks.
## Core Agents
### DroidAgent (Orchestrator)
Main coordinator that manages the workflow and routes between agents.
**Location**: `droidrun/agent/droid/droid_agent.py:75`
### ManagerAgent (Planner)
Creates high-level plans and monitors progress. Only used in reasoning mode.
**Location**: `droidrun/agent/manager/manager_agent.py:46`
### ExecutorAgent (Actor)
Executes specific actions for each subgoal. Only used in reasoning mode.
**Location**: `droidrun/agent/executor/executor_agent.py:46`
### CodeActAgent (Direct Executor)
Generates and executes Python code directly. Used in direct mode.
**Location**: `droidrun/agent/codeact/codeact_agent.py`
### ScripterAgent (Off-Device)
Handles Python computations without device interaction (API calls, calculations).
**Location**: `droidrun/agent/scripter/`
## Workflow Comparison
### Reasoning Mode Flow
```
Goal → Manager (creates plan) → Executor (executes action) →
Manager (checks result) → Executor (next action) → ...
```
### Direct Mode Flow
```
Goal → CodeActAgent (generates code) → Execute → Done
```
## Per-Agent Configuration
Configure different LLMs for each agent:
```yaml
# config.yaml
llm_profiles:
manager:
provider: Anthropic
model: claude-sonnet-4
executor:
provider: OpenAI
model: gpt-4o
codeact:
provider: GoogleGenAI
model: models/gemini-2.0-flash-exp
agent:
reasoning: true
max_steps: 15
```
## When to Use Each Mode
**Use Reasoning Mode for:**
- Multi-step tasks (booking flights, configuring settings)
- Tasks requiring planning and adaptation
- Complex workflows across multiple apps
**Use Direct Mode for:**
- Simple actions (screenshots, sending messages)
- Fast execution without planning overhead
- Well-defined single-step tasks
## Key Features
### Shared State
All agents share `DroidAgentState` for coordination:
- Action history
- Error tracking
- Memory and context
- Script results
## Dive Deeper
- [Agent Architecture](./agent-architecture) - Detailed design
- [Reasoning Mode](./reasoning-mode) - Manager/Executor workflow
- [Direct Mode](./direct-mode) - CodeAct workflow
- [Configuration](./configuration) - Setup and LLM profiles
- [Custom Tools](./custom-tools) - Extending functionality
- [Migration Guide](./migration-guide) - Upgrading from v3
+332
View File
@@ -0,0 +1,332 @@
---
title: 'Prompt Templates'
description: 'Customizing agent behavior with Jinja2 prompt templates.'
---
## Overview
Droidrun uses **Jinja2 templates** for agent prompts. You can customize agent behavior by passing custom template strings to `DroidAgent`:
```python
custom_prompts = {
"manager_system": "Your Jinja2 template here...",
"executor_system": "Another template...",
"codeact_system": "...",
"codeact_user": "...",
"scripter_system": "..."
}
agent = DroidAgent(
goal="Send an email",
config=config,
prompts=custom_prompts # Pass template strings, not file paths
)
```
**Important**: The `prompts` parameter accepts Jinja2 **template strings**, not file paths.
## How It Works
1. `DroidAgent` creates a `PromptResolver` with your custom prompts
2. Each agent checks if you provided a custom template for its key (e.g., "manager_system")
3. If found: uses your custom template
4. If not found: loads the default template from Droidrun's built-in files
5. Templates are rendered with context variables specific to each agent
## Available Prompt Keys
| Key | Agent | When Used |
|-----|-------|-----------|
| `manager_system` | Manager | Planning and reasoning (only in reasoning mode) |
| `executor_system` | Executor | Action selection (only in reasoning mode) |
| `codeact_system` | CodeAct | Direct execution (always used) |
| `codeact_user` | CodeAct | Task input formatting (always used) |
| `scripter_system` | Scripter | Off-device Python execution (when enabled) |
## Context Variables
Each agent has access to different variables in its templates:
### Manager
- `instruction` - User's goal
- `device_date` - Current device date/time
- `app_card` - App-specific guidance (empty if none available)
- `error_history` - List of recent failed actions with details
- `custom_tools_descriptions` - Custom tool documentation
- `scripter_execution_enabled` - Whether Scripter is available
- `available_secrets` - Available credential IDs
- `variables` - Custom variables passed to DroidAgent
- `output_schema` - Pydantic model schema (if provided)
### Executor
- `instruction` - User's goal
- `device_state` - Current UI tree
- `subgoal` - Current subgoal from Manager
- `atomic_actions` - Available actions (includes custom tools)
- `action_history` - Recent actions with outcomes
### CodeAct
**System prompt:**
- `tool_descriptions` - Available tool signatures
- `available_secrets` - Credential IDs
- `variables` - Custom variables
- `output_schema` - Output model schema (if provided)
**User prompt:**
- `goal` - Task description
- `variables` - Custom variables
### Scripter
- `task` - Task description from Manager
- `available_secrets` - Credential IDs
- `variables` - Custom variables
## Example: Custom Manager Prompt
```python
custom_prompts = {
"manager_system": """
You are a mobile automation planning agent.
Task: {{ instruction }}
Date: {{ device_date }}
{% if app_card %}
App guidance:
{{ app_card }}
{% endif %}
{% if error_history %}
Recent errors (you may be stuck):
{% for error in error_history %}
- Action: {{ error.action }}
Error: {{ error.error }}
{% endfor %}
{% endif %}
{% if custom_tools_descriptions %}
Custom tools:
{{ custom_tools_descriptions }}
{% endif %}
{% if variables.domain %}
Domain: {{ variables.domain }}
{% endif %}
Output format:
<thought>Your reasoning</thought>
<plan>
1. First step
2. Second step
3. DONE
</plan>
Or if complete:
<request_accomplished>
Task is done. Answer: ...
</request_accomplished>
"""
}
agent = DroidAgent(
goal="Send an email",
config=config,
prompts=custom_prompts,
variables={"domain": "finance"}
)
```
## Example: Using Custom Variables
Custom variables let you inject dynamic context into prompts:
```python
custom_prompts = {
"manager_system": """
Task: {{ instruction }}
{% if variables.budget %}
Budget limit: ${{ variables.budget }}
{% endif %}
{% if variables.priority %}
Priority: {{ variables.priority }}
{% endif %}
Guidelines:
{% for rule in variables.rules %}
- {{ rule }}
{% endfor %}
"""
}
agent = DroidAgent(
goal="Buy a phone",
config=config,
prompts=custom_prompts,
variables={
"budget": 1000,
"priority": "high",
"rules": ["Check reviews", "Compare prices", "Use coupons"]
}
)
```
## Jinja2 Syntax Reference
### Variables
```jinja2
{{ instruction }}
{{ variables.my_var }}
```
### Conditionals
```jinja2
{% if app_card %}
<app_card>{{ app_card }}</app_card>
{% endif %}
{% if error_history %}
You have {{ error_history | length }} errors
{% endif %}
```
### Loops
```jinja2
{% for error in error_history %}
- {{ error.action }}: {{ error.error }}
{% endfor %}
```
### Filters
```jinja2
{{ instruction | upper }}
{{ available_secrets | join(', ') }}
{{ error_history | length }}
```
## Best Practices
### 1. Use Clear Structure
```jinja2
<instruction>
{{ instruction }}
</instruction>
<guidelines>
1. Rule one
2. Rule two
</guidelines>
<output_format>
Expected format
</output_format>
```
### 2. Handle Missing Data with Conditionals
```jinja2
{% if app_card %}
<app_card>{{ app_card }}</app_card>
{% else %}
<note>No app-specific guidance available</note>
{% endif %}
```
### 3. Document Expected Variables
```jinja2
{# Expected variables:
- instruction: str - User's goal
- device_date: str - Current date/time
- app_card: str - App guidance (may be empty)
#}
```
### 4. Use Variables for Dynamic Behavior
```jinja2
{% if variables.strict_mode %}
<strict>
Follow instructions exactly. Do not make assumptions.
</strict>
{% endif %}
```
## Complete Example
```python
from droidrun import DroidAgent
from droidrun.config_manager import DroidrunConfig
# E-commerce automation with custom prompts
ecommerce_prompts = {
"manager_system": """
You are an e-commerce automation specialist.
Task: {{ instruction }}
Budget: ${{ variables.budget }}
{% if app_card %}
App info:
{{ app_card }}
{% endif %}
{% if error_history %}
Errors encountered:
{% for error in error_history %}
- {{ error.summary }}: {{ error.error }}
{% endfor %}
Consider changing your approach.
{% endif %}
Rules:
1. Verify product names exactly
2. Check prices before purchasing
3. Store order confirmations in memory
4. Never exceed budget
Output:
<thought>Your reasoning</thought>
<plan>
1. Step
2. DONE
</plan>
"""
}
config = DroidrunConfig()
agent = DroidAgent(
goal="Buy iPhone 15 Pro from Amazon",
config=config,
prompts=ecommerce_prompts,
variables={"budget": 1200}
)
result = await agent.run()
```
## Key Points
- Pass Jinja2 template **strings** (not file paths) to `DroidAgent(prompts={...})`
- Each agent has different available variables in its template
- Use `variables` parameter to inject custom context
- Templates are rendered at runtime with current state
- If no custom prompt provided, default templates are used
- Supports full Jinja2 syntax (conditionals, loops, filters)
## Related Topics
- [Agent Architecture](./agent-architecture) - How agents use prompts
- [Custom Tools](./custom-tools) - Adding tools referenced in prompts
- [App Cards](./app-cards) - Dynamic app-specific context
- [Configuration](./configuration) - Default prompt file paths
+165
View File
@@ -0,0 +1,165 @@
---
title: 'ScripterAgent'
description: 'Off-device Python execution for API calls, file operations, and data processing.'
---
## What is ScripterAgent?
**ScripterAgent** executes **off-device Python code** for tasks that don't require device interaction. It's triggered by ManagerAgent when API calls, file operations, or data processing are needed.
ScripterAgent enables:
- **API calls**: REST APIs, webhooks, database queries
- **File operations**: Reading, writing, parsing files
- **Data processing**: JSON/CSV parsing, transformations, filtering
**Key difference**: ScripterAgent runs code that **doesn't interact with the device**, while CodeActAgent generates code that **uses device atomic actions**.
## How It Works
ManagerAgent delegates tasks to ScripterAgent by providing **full context and a high-level task description**. ScripterAgent is a **ReAct agent** that follows a think-execute-observe loop:
1. **Receives task** from Manager with full context
2. **Thinks** and generates Python code to make progress
3. **Executes** the code and observes the output
4. **Repeats** steps 2-3 until task is complete
5. **Returns message** to Manager summarizing the results
ScripterAgent signals completion by returning a message **without code** (not a function call). Variables persist across iterations like a Jupyter notebook.
## Examples
### API Calls
**Manager delegates**:
```
User needs current weather in San Francisco to decide clothing.
Task: Fetch weather from API and extract temperature + conditions
API: https://api.weather.com/forecast?city=San Francisco
```
**ScripterAgent (Iteration 1)**:
```python
# Think: I need to make an API request to get weather data
import requests
response = requests.get(
"https://api.weather.com/forecast",
params={"city": "San Francisco"}
)
print(response.json())
```
**Output**: `{'temp': 62, 'description': 'Partly cloudy', 'humidity': 65}`
**ScripterAgent (Final)**:
```
The weather in San Francisco is 62°F with partly cloudy conditions.
```
*(Manager receives this message)*
### Webhook Notifications
**Manager delegates**:
```
Flight booking was completed successfully at 2024-01-15T10:30:00.
Task: Send status update to monitoring webhook
Webhook URL: https://webhook.site/your-webhook-id
Include: task name, status, timestamp
```
**ScripterAgent**:
```python
import requests
payload = {
"task": "Flight booking completed",
"status": "success",
"timestamp": "2024-01-15T10:30:00"
}
response = requests.post(
"https://webhook.site/your-webhook-id",
json=payload
)
print(f"Webhook response: {response.status_code}")
```
**Output**: `Webhook response: 200`
**ScripterAgent (Final)**:
```
Successfully sent webhook notification. Server responded with status 200.
```
### File Operations
**Manager delegates**:
```
Contacts file at /tmp/contacts.csv contains all contacts.
Task: Filter contacts from New York and save as JSON
Output: /tmp/filtered_contacts.json
```
**ScripterAgent (Iteration 1)**:
```python
import csv
with open('/tmp/contacts.csv', 'r') as f:
reader = csv.DictReader(f)
contacts = list(reader)
print(f"Total contacts: {len(contacts)}")
print(f"Sample: {contacts[0]}")
```
**Output**: `Total contacts: 150` `Sample: {'name': 'John', 'city': 'New York', 'email': '...'}`
**ScripterAgent (Iteration 2)**:
```python
import json
# Filter New York contacts
ny_contacts = [c for c in contacts if c['city'] == 'New York']
with open('/tmp/filtered_contacts.json', 'w') as f:
json.dump(ny_contacts, f, indent=2)
print(f"Filtered {len(ny_contacts)} contacts from New York")
```
**Output**: `Filtered 23 contacts from New York`
**ScripterAgent (Final)**:
```
Found 23 contacts from New York and saved them to /tmp/filtered_contacts.json
```
## Configuration
```yaml
agent:
scripter:
max_steps: 10
safe_execution:
enabled: true
allowed_modules:
- requests # API calls, webhooks
- json # JSON parsing
- csv # CSV file operations
- pathlib # File path handling
blocked_modules:
- subprocess # Process spawning
- os # System operations
- sys # System modifications
```
**Safe execution** restricts dangerous operations by default. Only modules in `allowed_modules` can be imported.
## Key Points
- **ReAct agent**: Think-execute-observe loop until task complete
- **Off-device only**: No device interactions
- **State persistence**: Variables persist across iterations (Jupyter-style)
- **Completion signal**: Returns message without code when done
- **Safe by default**: Restricted imports/builtins
## Related Topics
- [Multi-Agent Architecture](./agent-architecture) - ScripterAgent in the workflow
- [Reasoning Mode](./reasoning-mode) - How Manager delegates to ScripterAgent
- [Safe Execution](./safe-execution) - Security restrictions
+60
View File
@@ -0,0 +1,60 @@
---
title: 'Shared State'
description: 'DroidAgentState - the coordination mechanism for multi-agent workflow communication.'
---
## What is Shared State?
**DroidAgentState** is a Pydantic model that serves as the **central coordination mechanism** for Droidrun v4's multi-agent workflow. It's a shared data structure that all agents (Manager, Executor, CodeAct, Scripter) can read from and write to.
Shared state enables:
- **Cross-agent communication**: Agents share information about actions, results, and errors
- **Progress tracking**: Step counts, action history, visited apps/screens
- **Memory management**: Agent memory, custom variables, user session data
- **Error coordination**: Error flags, escalation thresholds, error descriptions
**Key insight**: Shared state replaces complex message passing. Instead of sending data back and forth, agents update a single shared object.
## Core State Fields
```python
class DroidAgentState(BaseModel):
# Task context
instruction: str = "" # Original task
step_number: int = 0 # Current step
# Device state
formatted_device_state: str = "" # Human-readable state
current_package_name: str = "" # Current app
current_activity_name: str = "" # Current screen
# Action tracking
action_history: List[Dict] = [] # All actions taken
action_outcomes: List[bool] = [] # Success/failure
summary_history: List[str] = [] # Action summaries
# Memory
memory: str = "" # Agent persistent memory
# Planning (Manager)
plan: str = "" # Current plan
current_subgoal: str = "" # Current subgoal
manager_answer: str = "" # Answer-type responses
# Error handling
error_flag_plan: bool = False # Signal error to Manager
error_descriptions: List[str] = [] # Error messages
# Script execution (Scripter)
scripter_history: List[Dict] = [] # Scripter results
last_scripter_success: bool = True # Last execution status
# Custom variables
custom_variables: Dict = {} # User-defined data
```
## Related Topics
- [Agent Architecture](./agent-architecture) - How agents use shared state
- [Reasoning Mode](./reasoning-mode) - Manager/Executor state updates
- [ScripterAgent](./scripter-agent) - Scripter result storage
File diff suppressed because it is too large Load Diff
+12 -12
View File
@@ -5,7 +5,7 @@ description: 'Enhance agent performance with app-specific guidance and best prac
# App-Specific Instruction Cards
DroidRun's app card system provides app-specific guidance to agents, helping them understand how to operate specific apps more effectively. App cards are automatically loaded based on the currently active app and injected into agent prompts to improve task success rates.
Droidrun's app card system provides app-specific guidance to agents, helping them understand how to operate specific apps more effectively. App cards are automatically loaded based on the currently active app and injected into agent prompts to improve task success rates.
## Overview
@@ -17,7 +17,7 @@ App cards are markdown documents that contain:
- **Known gotchas**: Quirks, issues, or important notes about the app
- **Search syntax**: Special search operators or filters (for apps like Gmail, file managers, etc.)
When a DroidRun agent detects it's operating within a specific app (by package name), it automatically loads the corresponding app card and uses that knowledge to make better decisions.
When a Droidrun agent detects it's operating within a specific app (by package name), it automatically loads the corresponding app card and uses that knowledge to make better decisions.
---
@@ -25,8 +25,8 @@ When a DroidRun agent detects it's operating within a specific app (by package n
### Automatic Detection and Loading
1. **Package Detection**: DroidRun monitors the current foreground app package name (e.g., "com.google.android.gm" for Gmail)
2. **Async Loading**: When the package or instruction changes, DroidRun starts loading the app card in the background
1. **Package Detection**: Droidrun monitors the current foreground app package name (e.g., "com.google.android.gm" for Gmail)
2. **Async Loading**: When the package or instruction changes, Droidrun starts loading the app card in the background
3. **Caching**: All providers use in-memory caching with `(package_name, instruction)` as cache keys
4. **Prompt Injection**: The app card content is injected into the Manager's system prompt within `<app_card>` tags
5. **Context-Aware Actions**: The agent uses the app-specific guidance to perform actions more effectively
@@ -53,7 +53,7 @@ App cards are used by the **Manager Agent** in reasoning mode (`--reasoning` fla
## Built-in App Cards
DroidRun ships with app cards for popular apps. Currently included:
Droidrun ships with app cards for popular apps. Currently included:
| App | Package Name | Features Covered |
|-----|--------------|------------------|
@@ -188,9 +188,9 @@ adb shell pm list packages | grep keyword
adb shell dumpsys window windows | grep -E 'mCurrentFocus'
```
**Using DroidRun Debug Mode:**
**Using Droidrun Debug Mode:**
```sh
# Run DroidRun with debug logging
# Run Droidrun with debug logging
droidrun run "Open the app" --debug
# The package name will appear in logs when the app is opened
@@ -230,7 +230,7 @@ Add an entry to `app_cards.json`:
### Step 4: Test the App Card
Run DroidRun with the app and verify the app card is loaded:
Run Droidrun with the app and verify the app card is loaded:
```sh
droidrun run "Open myapp and do something" --reasoning --debug
@@ -391,10 +391,10 @@ App cards are cached in memory. Restart the agent or call `provider.clear_cache(
## Related Documentation
- [CLI Usage](/docs/v4/guides/cli) - DroidRun CLI command reference
- [Configuration](/docs/v4/concepts/configuration) - Configuration system details
- [Agent Architecture](/docs/v4/concepts/architecture) - How agents use app cards
- [Manager Agent](/docs/v4/sdk/droid-agent#manager-agent) - Agent that uses app cards
- [CLI Usage](/v4/guides/cli) - Droidrun CLI command reference
- [Configuration](/v4/sdk/configuration) - Configuration system details
- [Agent Architecture](/v4/concepts/agent-architecture) - How agents use app cards
- [Manager Agent](/v4/sdk/droid-agent#manager-agent) - Agent that uses app cards
---
+264 -937
View File
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+39 -49
View File
@@ -1,15 +1,15 @@
---
title: 'Custom Variables'
description: 'Pass dynamic data and configuration to your DroidRun agents'
description: 'Pass dynamic data and configuration to your Droidrun agents'
---
# Custom Variables
Pass dynamic data to your DroidRun agents using the `variables` parameter. Variables enable parameterized workflows and reusable automation.
Pass dynamic data to your Droidrun agents using the `variables` parameter. Variables enable parameterized workflows and reusable automation.
<Warning>
**Important Limitation**: Custom variables are **only accessible via agent prompts**, not directly to custom tools. Tools don't have access to `shared_state.custom_variables`. See [workaround below](#accessing-variables-in-custom-tools).
</Warning>
Custom variables are accessible in:
- **Agent prompts** via custom Jinja2 templates
- **Custom tools** via `shared_state.custom_variables`
---
@@ -17,8 +17,7 @@ Pass dynamic data to your DroidRun agents using the `variables` parameter. Varia
```python
from droidrun.agent.droid import DroidAgent
from droidrun.tools import AdbTools
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Define custom prompts that render variables
custom_prompts = {
@@ -35,12 +34,11 @@ Available variables:
}
# Create agent with variables
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Send email to recipient with subject",
config=config,
tools=AdbTools(),
variables={"recipient": "john@example.com", "subject": "Update"},
prompts=custom_prompts # Required to see variables
)
@@ -64,7 +62,7 @@ When you pass `variables` to `DroidAgent`:
**Access Summary:**
- ✅ Agent prompts (via custom Jinja2 templates)
- ✅ Agents can read and pass to tools as arguments
- ❌ NOT directly accessible to custom tools
- ✅ Custom tools (via `shared_state.custom_variables`)
---
@@ -72,13 +70,12 @@ When you pass `variables` to `DroidAgent`:
```python
from droidrun.agent.droid import DroidAgent
from droidrun.tools import AdbTools
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Define variables
variables = {
"recipient": "alice@example.com",
"message": "Hello from DroidRun!"
"message": "Hello from Droidrun!"
}
# Custom prompt to render variables
@@ -98,12 +95,11 @@ Use these variables when executing tasks.
}
# Create agent
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Send message to recipient",
config=config,
tools=AdbTools(),
variables=variables,
prompts=custom_prompts
)
@@ -125,51 +121,45 @@ Customize these prompts to render variables:
## Accessing Variables in Custom Tools
Custom tools **cannot** directly access `DroidAgentState.custom_variables` because they don't have access to the shared state.
### Workaround: Pass as Tool Arguments (Recommended)
Design tools to accept parameters, then let the agent pass variable values:
Custom tools can directly access variables via the `shared_state` keyword argument:
```python
from droidrun.agent.droid import DroidAgent
from droidrun.tools import AdbTools
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
def send_notification(tools_instance, title: str, channel: str):
"""Send a notification to a specific channel."""
def send_notification(title: str, *, tools=None, shared_state=None, **kwargs):
"""Send a notification using channel from custom variables.
Args:
title: Notification title
tools: Tools instance (optional, injected automatically)
shared_state: DroidAgentState (optional, injected automatically)
"""
if not shared_state:
return "Error: shared_state required"
# Access custom variables
channel = shared_state.custom_variables.get("notification_channel", "default")
return f"Sent '{title}' to {channel}"
custom_tools = {
"send_notification": {
"arguments": ["title", "channel"],
"description": "Send a notification with title and channel",
"arguments": ["title"],
"description": "Send a notification with title. Usage: {\"action\": \"send_notification\", \"title\": \"Alert\"}",
"function": send_notification
}
}
custom_prompts = {
"codeact_system": """
{% if variables %}
Variables: {% for k, v in variables.items() %}{{ k }}={{ v }} {% endfor %}
{% endif %}
"""
}
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Send notification to alerts channel",
goal="Send notification with title 'Alert'",
config=config,
tools=AdbTools(),
custom_tools=custom_tools,
variables={"notification_channel": "alerts"},
prompts=custom_prompts
variables={"notification_channel": "alerts"}
)
```
The agent sees `notification_channel="alerts"` in the prompt and calls `send_notification(title="...", channel="alerts")`.
---
## Use Cases
@@ -208,14 +198,14 @@ agent = DroidAgent(
## Key Points
1. **Custom prompts required** - Default prompts don't render variables
2. **Access via prompts only** - Tools cannot access `shared_state.custom_variables` directly
3. **Workaround** - Agents read variables from prompts → pass as tool arguments
4. **Available to all agents** - Manager, Executor, CodeAct, Scripter all receive variables
5. **Jinja2 templates** - Use `{% if variables %}` blocks in custom prompts
1. **Custom prompts required** - Default prompts don't render variables in agent context
2. **Direct access in tools** - Custom tools access `shared_state.custom_variables` via keyword argument
3. **Available to all agents** - Manager, Executor, CodeAct, Scripter all receive variables
4. **Jinja2 templates** - Use `{% if variables %}` blocks in custom prompts
5. **Auto-injection** - `tools` and `shared_state` are injected automatically by Droidrun
## Related Documentation
- [Custom Prompts](/docs/v4/guides/prompts) - How to customize agent prompts
- [Custom Tools](/docs/v4/guides/custom-tools) - Creating custom tool functions
- [DroidAgent SDK](/docs/v4/sdk/droid-agent) - Complete API reference
- [Custom Prompts](/v4/concepts/prompts) - How to customize agent prompts
- [Custom Tools and Credentials](/v4/guides/custom-tools-credentials) - Creating custom tool functions
- [DroidAgent SDK](/v4/sdk/droid-agent) - Complete API reference
File diff suppressed because it is too large Load Diff
+4 -5
View File
@@ -2,7 +2,7 @@
title: Guides Overview
---
Welcome to the DroidRun v4 Guides! This section provides step-by-step instructions and best practices for using DroidRun. Each guide focuses on a specific aspect of the framework, from device setup to advanced automation patterns.
Welcome to the Droidrun v4 Guides! This section provides step-by-step instructions and best practices for using Droidrun. Each guide focuses on a specific aspect of the framework, from device setup to advanced automation patterns.
---
@@ -70,7 +70,7 @@ Welcome to the DroidRun v4 Guides! This section provides step-by-step instructio
## Quick Start Paths
### New to DroidRun?
### New to Droidrun?
1. [CLI Reference](./cli) - Learn CLI commands and usage
2. [Device Setup](./device-setup) - Get your device ready
3. [Configuration System](./configuration) - Configure agents and LLMs
@@ -81,7 +81,7 @@ Welcome to the DroidRun v4 Guides! This section provides step-by-step instructio
2. [Custom Variables](./custom-variables) - Make workflows reusable
3. [App Cards](./app-cards) - Improve success rates for specific apps
### Extending DroidRun?
### Extending Droidrun?
1. [Custom Tools & Credentials](./custom-tools-credentials) - Add new capabilities
2. [Configuration System](./configuration) - Advanced customization
3. Check [SDK Reference](../sdk/droid-agent) for programmatic usage
@@ -92,8 +92,7 @@ Welcome to the DroidRun v4 Guides! This section provides step-by-step instructio
**Core Concepts:**
- [Architecture](../concepts/architecture) - Multi-agent system overview
- [Workflow Architecture](../concepts/workflow-architecture) - Event-driven coordination
- [Event Streaming](../concepts/event-streaming) - Real-time monitoring
- [Events and Workflows](../concepts/events-and-workflows) - Event-driven coordination and real-time monitoring
**SDK Reference:**
- [DroidAgent](../sdk/droid-agent) - Main agent class
File diff suppressed because it is too large Load Diff
+19 -34
View File
@@ -5,7 +5,7 @@ description: 'Configure anonymous telemetry, Phoenix tracing, and trajectory rec
# Telemetry & Tracing
DroidRun provides three monitoring capabilities:
Droidrun provides three monitoring capabilities:
1. **Anonymous Telemetry** - Usage analytics sent to PostHog (enabled by default)
2. **Arize Phoenix Tracing** - Real-time LLM and agent execution tracing (opt-in)
@@ -46,7 +46,7 @@ phoenix serve
The server starts at `http://localhost:6006` and provides a web UI for viewing traces.
**3. Enable tracing in DroidRun:**
**3. Enable tracing in Droidrun:**
**Via CLI:**
@@ -65,9 +65,9 @@ tracing:
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
config = DroidRunConfig()
config = DroidrunConfig()
config.tracing.enabled = True
agent = DroidAgent(goal="Open settings", config=config)
@@ -104,37 +104,22 @@ Environment variable names are lowercase: `phoenix_url` and `phoenix_project_nam
## Anonymous Telemetry
DroidRun collects anonymous usage analytics via PostHog to help improve the framework. This is completely separate from Phoenix tracing and trajectory recording.
**Telemetry is enabled by default.** The framework tracks agent initialization, task completion, and usage patterns to help prioritize features and improve reliability.
Droidrun collects anonymous usage analytics via PostHog to help improve the framework. Telemetry is enabled by default.
### What's Collected
**Agent initialization:**
- LLM provider and model names (e.g., "GoogleGenAI", "models/gemini-2.5-pro")
- Configuration settings (max steps, timeout, vision/reasoning mode, debug flags)
- Tool platform (e.g., "AdbTools", "IOSTools")
- Trajectory save level ("none", "step", "action")
- Run type (CLI, developer, web)
**During execution:**
- App packages and activities visited (e.g., "com.android.settings")
- Step number when visiting new packages
**Task completion:**
- Success/failure status and reason
- Total steps taken
- Count of unique packages and activities visited
**User identifier:**
- Task goals and completion status
- LLM providers and configuration settings (max steps, timeout, vision/reasoning mode, debug flags)
- Available tools and packages/activities visited
- Anonymous UUID stored in `~/.droidrun/user_id`
- No personal information attached
**What's NOT collected:**
### What's NOT Collected
- Screenshots or screen content
- User input or task goals (the actual text you type)
- API keys or credentials
- Device serial numbers or personal identifiers
- LLM responses
- Device serial numbers
- Detailed action history
### Disable Telemetry
@@ -190,9 +175,9 @@ logging:
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
config = DroidRunConfig()
config = DroidrunConfig()
config.logging.save_trajectory = "action"
agent = DroidAgent(goal="Open settings", config=config)
@@ -232,7 +217,7 @@ Use these files to:
| Feature | Data Location | Default | Purpose | Opt-In/Out |
|---------|--------------|---------|---------|------------|
| **Anonymous Telemetry** | PostHog (cloud) | Enabled | Help improve DroidRun | `DROIDRUN_TELEMETRY_ENABLED=false` |
| **Anonymous Telemetry** | PostHog (cloud) | Enabled | Help improve Droidrun | `DROIDRUN_TELEMETRY_ENABLED=false` |
| **Phoenix Tracing** | Phoenix server (local/cloud) | Disabled | Debug LLM calls and agent flow | `--tracing` flag or config |
| **Trajectory Recording** | Local disk (`trajectories/`) | Disabled | Offline debugging with screenshots | `--save-trajectory step/action` |
@@ -240,6 +225,6 @@ Use these files to:
## Related Documentation
- [Configuration System](/docs/v4/guides/configuration) - Configure tracing and telemetry settings
- [Event Streaming](/docs/v4/concepts/event-streaming) - Build custom monitoring integrations
- [CLI Usage](/docs/v4/guides/cli) - Command-line flags for monitoring
- [Configuration System](/v4/sdk/configuration) - Configure tracing and telemetry settings
- [Events and Workflows](/v4/concepts/events-and-workflows) - Build custom monitoring integrations
- [CLI Usage](/v4/guides/cli) - Command-line flags for monitoring
+33 -141
View File
@@ -1,22 +1,22 @@
---
title: 'Overview'
description: 'DroidRun is a powerful framework that enables you to control Android and iOS devices through intelligent LLM agents. Build sophisticated mobile automation workflows with natural language commands.'
description: 'Droidrun is a powerful framework that enables you to control Android and iOS devices through intelligent LLM agents. Build sophisticated mobile automation workflows with natural language commands.'
---
## What is DroidRun?
## What is Droidrun?
DroidRun empowers developers to automate mobile device interactions using AI-powered agents. Whether you're building testing frameworks, automating data collection, or creating intelligent mobile workflows, DroidRun provides the flexibility and power you need.
Droidrun empowers developers to automate mobile device interactions using AI-powered agents. Whether you're building testing frameworks, automating data collection, or creating intelligent mobile workflows, Droidrun provides the flexibility and power you need.
Built on [LlamaIndex workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/), DroidRun features a sophisticated multi-agent architecture that can handle everything from simple atomic tasks to complex multi-step workflows requiring strategic planning and error recovery.
Built on [LlamaIndex workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/), Droidrun features a sophisticated multi-agent architecture that can handle everything from simple atomic tasks to complex multi-step workflows requiring strategic planning and error recovery.
<CardGroup cols={2}>
<Card title="Quickstart" icon="rocket" href="/v4/quickstart">
Get up and running with DroidRun in minutes
Get up and running with Droidrun in minutes
</Card>
<Card title="Multi-Agent Architecture" icon="sitemap" href="/v4/concepts/architecture">
<Card title="Multi-Agent Architecture" icon="sitemap" href="/v4/concepts/agent-architecture">
Understand the hierarchical agent system
</Card>
<Card title="Event Streaming" icon="stream" href="/v4/concepts/event-streaming">
<Card title="Event Streaming" icon="stream" href="/v4/concepts/events-and-workflows">
Monitor and debug with real-time events
</Card>
<Card title="SDK Reference" icon="code" href="/v4/sdk">
@@ -30,7 +30,7 @@ Built on [LlamaIndex workflows](https://docs.llamaindex.ai/en/stable/understandi
### Multi-Agent Architecture
DroidRun v4 introduces a hierarchical multi-agent system with specialized agents for different responsibilities:
Droidrun v4 introduces a hierarchical multi-agent system with specialized agents for different responsibilities:
- **DroidAgent** - Main coordinator orchestrating execution flow
- **ManagerAgent** - High-level planning and strategic decision making
@@ -42,21 +42,21 @@ DroidRun v4 introduces a hierarchical multi-agent system with specialized agents
- **Direct Mode** (`reasoning=False`) - Fast, single-agent execution for simple tasks
- **Reasoning Mode** (`reasoning=True`) - Strategic planning with Manager/Executor workflow for complex tasks
<Card title="Learn More" icon="book" href="/v4/concepts/architecture">
<Card title="Learn More" icon="book" href="/v4/concepts/agent-architecture">
Deep dive into the agent architecture
</Card>
### Event Streaming System
Monitor and debug agent execution in real-time with DroidRun's comprehensive event system:
Monitor and debug agent execution in real-time with Droidrun's comprehensive event system:
```python
from droidrun import DroidAgent, ResultEvent
from droidrun.agent.droid.events import ManagerPlanEvent, ExecutorResultEvent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Create config with defaults
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(goal="Open Settings and enable WiFi", config=config)
handler = agent.run()
@@ -77,7 +77,7 @@ Events provide rich metadata for:
- Debugging and trajectory analysis
- Integration with external systems
<Card title="Event Streaming Guide" icon="chart-line" href="/v4/concepts/event-streaming">
<Card title="Event Streaming Guide" icon="chart-line" href="/v4/concepts/events-and-workflows">
Master the event system
</Card>
@@ -116,7 +116,7 @@ llm_profiles:
- Ollama
- DeepSeek
<Card title="Configuration Guide" icon="sliders" href="/v4/guides/configuration">
<Card title="Configuration Guide" icon="sliders" href="/v4/sdk/configuration">
Configure your agents
</Card>
@@ -127,7 +127,7 @@ Extract type-safe, validated data from device interactions using Pydantic models
```python
from pydantic import BaseModel, Field
from droidrun import DroidAgent, ResultEvent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
class ContactInfo(BaseModel):
"""Contact information extracted from device."""
@@ -136,7 +136,7 @@ class ContactInfo(BaseModel):
email: str = Field(description="Email address")
# Create config with defaults
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Find John Smith's contact and extract details",
@@ -162,17 +162,18 @@ Extend agent capabilities with custom tools and secure credential management:
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
def search_database(query: str) -> str:
def search_database(query: str, **kwargs) -> str:
"""Search the local database."""
# Your implementation
return f"Results for: {query}"
# Your database search implementation
results = db.query(query) # Example
return f"Found {len(results)} results for '{query}'"
custom_tools = {
"search_database": {
"signature": "search_database(query: str) -> str",
"description": "Search the local database for information",
"arguments": ["query"],
"description": "Search local database. Usage: {\"action\": \"search_database\", \"query\": \"search term\"}",
"function": search_database
}
}
@@ -184,7 +185,7 @@ credentials = {
}
# Create config with defaults
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Search database and email results",
@@ -253,7 +254,7 @@ Inject dynamic data into agent prompts:
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Define custom prompts that render variables
custom_prompts = {
@@ -270,7 +271,7 @@ Available variables:
}
# Create config with defaults
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Complete task using context",
@@ -287,7 +288,7 @@ agent = DroidAgent(
Variables are accessible in:
- ✅ Agent prompts (via custom Jinja2 templates)
- ✅ Agents can read and pass to tools as arguments
- ❌ NOT directly accessible to custom tools
- ✅ Accessible to custom tools
<Card title="Custom Variables Guide" icon="brackets-curly" href="/v4/guides/custom-variables">
Use dynamic variables
@@ -297,7 +298,7 @@ Variables are accessible in:
## Two Deployment Options
DroidRun offers flexible deployment to match your workflow:
Droidrun offers flexible deployment to match your workflow:
<CardGroup cols={2}>
<Card icon="mobile" title="Physical Device" href="/v4/quickstart">
@@ -333,128 +334,19 @@ DroidRun offers flexible deployment to match your workflow:
---
## What's New in v4
DroidRun v4 is a major architectural upgrade from v3, introducing powerful new capabilities:
### New Agent System
- **ScripterAgent** - Off-device Python code execution for computations, API calls, and data processing
- **Structured Output** - Automatic extraction of typed data using Pydantic models
- **Helper Workflows** - App opener and text manipulator for specialized tasks
- **Multi-agent coordination** - Hierarchical workflow orchestration with shared state
### Event Streaming
- **Real-time event system** - Stream execution events for monitoring and debugging
- **Rich event types** - Workflow events (Manager, Executor, Scripter), action events, and state events
- **Custom handlers** - Build custom UIs, webhooks, and integrations
- **Trajectory recording** - Capture screenshots, UI states, and action history
### LLM Configuration
- **Per-agent LLM profiles** - Different models for Manager, Executor, CodeAct, etc.
- **Profile system** - Reusable LLM configurations in YAML
- **Provider flexibility** - Mix and match providers (OpenAI, Anthropic, Google, etc.)
- **Cost optimization** - Use powerful models for planning, fast models for actions
### Enhanced Features
- **Custom variables** - Store and access custom data throughout execution
- **Structured output** - Type-safe data extraction with Pydantic models
- **Safe execution** - Sandboxed code execution with import/builtin restrictions
- **Credentials** - Auto-injection as custom tools with direct dict or file-based config
- **App cards** - App-specific guidance (local, server, or composite modes)
- **Improved error handling** - Error escalation and plan adjustments
### Developer Experience
- **DroidRunConfig** - Clean configuration API with `from_yaml()` loading
- **Improved logging** - Rich console output with progress tracking
- **Tracing integration** - Arize Phoenix support for execution tracing
- **Better type hints** - Complete type annotations and IDE support
- **Event-driven architecture** - LlamaIndex workflows for easy integration
---
## Quick Start
### Installation
```bash
# Install for cli usage.
uv tool install 'droidrun[google,anthropic,openai,deepseek,ollama,openrouter]'
# Or for developing on top of Droidrun.
uv pip install 'droidrun[google,anthropic,openai,deepseek,ollama,openrouter]'
```
### Basic Usage
```python
import asyncio
from droidrun import DroidAgent, ResultEvent
from droidrun.config_manager.config_manager import DroidRunConfig
async def main():
# Create config with defaults
config = DroidRunConfig()
config.agent.max_steps = 20
config.agent.reasoning = True
# Or load configuration from YAML
# config = DroidRunConfig.from_yaml("config.yaml")
# Create agent (tools auto-created from device config)
agent = DroidAgent(
goal="Open Settings and enable WiFi",
config=config
)
# Run agent and get result
handler = agent.run()
result: ResultEvent = await handler
if result.success:
print(f"Success: {result.reason}")
else:
print(f"Failed: {result.reason}")
asyncio.run(main())
```
### CLI Usage
```bash
# Run a command
droidrun run "Open Settings and enable WiFi"
# With reasoning mode
droidrun run "Set up a new alarm for 7 AM" --reasoning
# With vision
droidrun run "Take a screenshot and describe what's on screen" --vision
# Device management
droidrun devices # List devices
droidrun setup # Install Portal APK
droidrun ping # Test connection
```
<Card title="Complete Quickstart" icon="rocket" href="/v4/quickstart">
Follow the full quickstart guide
</Card>
---
## Key Concepts
<CardGroup cols={2}>
<Card title="Agent Architecture" icon="sitemap" href="/v4/concepts/architecture">
<Card title="Agent Architecture" icon="sitemap" href="/v4/concepts/agent-architecture">
Multi-agent system with specialized roles
</Card>
<Card title="Workflow Events" icon="diagram-project" href="/v4/concepts/workflow-architecture">
<Card title="Workflow Events" icon="diagram-project" href="/v4/concepts/events-and-workflows">
LlamaIndex workflow integration
</Card>
<Card title="Event Streaming" icon="stream" href="/v4/concepts/event-streaming">
<Card title="Event Streaming" icon="stream" href="/v4/concepts/events-and-workflows">
Real-time monitoring and debugging
</Card>
<Card title="Configuration" icon="gear" href="/v4/guides/configuration">
<Card title="Configuration" icon="gear" href="/v4/sdk/configuration">
Agent, LLM, and system configuration
</Card>
</CardGroup>
@@ -499,7 +391,7 @@ droidrun ping # Test connection
Follow for updates and news
</Card>
<Card title="Benchmark" icon="chart-bar" href="https://droidrun.ai/benchmark">
See DroidRun's performance metrics
See Droidrun's performance metrics
</Card>
<Card title="Cloud Platform" icon="cloud" href="https://cloud.droidrun.ai">
Try the managed cloud service
+32 -33
View File
@@ -1,6 +1,6 @@
---
title: 'Quickstart'
description: 'Get up and running with DroidRun v4 quickly and effectively'
description: 'Get up and running with Droidrun v4 quickly and effectively'
---
<iframe
@@ -12,11 +12,11 @@ description: 'Get up and running with DroidRun v4 quickly and effectively'
allowFullScreen
></iframe>
This guide will help you get DroidRun v4 installed and running quickly, controlling your Android device through natural language in minutes. DroidRun v4 introduces a config-driven architecture, making it easier to customize agent behavior, LLM selection, and execution modes.
This guide will help you get Droidrun v4 installed and running quickly, controlling your Android device through natural language in minutes. Droidrun v4 introduces a config-driven architecture, making it easier to customize agent behavior, LLM selection, and execution modes.
### Prerequisites
Before installing DroidRun, ensure you have:
Before installing Droidrun, ensure you have:
1. **Python 3.11+** installed on your system
2. [Android Debug Bridge (adb)](https://developer.android.com/studio/releases/platform-tools) installed and configured
@@ -27,7 +27,7 @@ Before installing DroidRun, ensure you have:
### Installation
DroidRun v4 is installed using [`uv`](https://docs.astral.sh/uv/), a fast Python package installer and resolver.
Droidrun v4 is installed using [`uv`](https://docs.astral.sh/uv/), a fast Python package installer and resolver.
**Install uv (if not already installed):**
@@ -57,7 +57,7 @@ If you only need specific providers, you can install just those. For example, `u
### Setup the Portal APK
DroidRun requires the Portal app to be installed on your Android device for device control. The Portal app provides accessibility services that expose the UI accessibility tree, enabling the agent to see and interact with UI elements.
Droidrun requires the Portal app to be installed on your Android device for device control. The Portal app provides accessibility services that expose the UI accessibility tree, enabling the agent to see and interact with UI elements.
```bash
droidrun setup
@@ -70,7 +70,7 @@ This command automatically:
### Test Connection
Verify that DroidRun can communicate with your device:
Verify that Droidrun can communicate with your device:
```bash
droidrun ping
@@ -85,7 +85,7 @@ If successful, you'll see:
### Configure Your LLM
DroidRun v4 uses a configuration-driven approach. On first run, DroidRun creates a `config.yaml` file with default settings. You'll need to set your API key for your chosen LLM provider.
Droidrun v4 uses a configuration-driven approach. On first run, Droidrun creates a `config.yaml` file with default settings. You'll need to set your API key for your chosen LLM provider.
**Set your API key:**
@@ -136,11 +136,11 @@ For complex automation or integration into your Python projects, create a script
```python
import asyncio
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
async def main():
# Use default configuration with built-in LLM profiles
config = DroidRunConfig()
config = DroidrunConfig()
# Create agent
# LLMs are automatically loaded from config.llm_profiles
@@ -166,12 +166,12 @@ if __name__ == "__main__":
```python
import asyncio
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig, AgentConfig
from droidrun.config_manager.config_manager import DroidrunConfig, AgentConfig
from llama_index.llms.openai import OpenAI
async def main():
# Load base configuration
config = DroidRunConfig()
config = DroidrunConfig()
# Create custom agent config with overrides
agent_config = AgentConfig(
@@ -201,23 +201,23 @@ if __name__ == "__main__":
### Configuration Options
DroidRun v4 provides flexible configuration through:
Droidrun v4 provides flexible configuration through:
**1. Default Configuration (No File Required):**
```python
config = DroidRunConfig() # Uses built-in defaults
config = DroidrunConfig() # Uses built-in defaults
```
**2. Load from YAML File:**
```python
config = DroidRunConfig.from_yaml("config.yaml") # Load custom config
config = DroidrunConfig.from_yaml("config.yaml") # Load custom config
```
**3. Programmatic Configuration:**
```python
from droidrun.config_manager.config_manager import DroidRunConfig, AgentConfig
from droidrun.config_manager.config_manager import DroidrunConfig, AgentConfig
config = DroidRunConfig()
config = DroidrunConfig()
config.agent.max_steps = 25
config.agent.reasoning = True
config.agent.codeact.vision = True
@@ -233,18 +233,18 @@ config.agent.codeact.vision = True
<Tip>
For detailed configuration options, see the [Configuration Guide](/v4/guides/configuration). The CLI uses `ConfigManager` which auto-creates config.yaml on first run.
For detailed configuration options, see the [Configuration Guide](/v4/sdk/configuration). The CLI uses `ConfigManager` which auto-creates config.yaml on first run.
</Tip>
### Structured Output Extraction
DroidRun v4 supports structured output extraction using Pydantic models:
Droidrun v4 supports structured output extraction using Pydantic models:
```python
import asyncio
from pydantic import BaseModel, Field
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
class BatteryInfo(BaseModel):
"""Battery information."""
@@ -253,7 +253,7 @@ class BatteryInfo(BaseModel):
temperature: float = Field(description="Battery temperature in Celsius")
async def main():
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Check battery status",
@@ -276,7 +276,7 @@ if __name__ == "__main__":
### Execution Modes
DroidRun v4 supports two execution modes:
Droidrun v4 supports two execution modes:
**Direct Mode (Default):**
- Fast execution for simple tasks
@@ -290,9 +290,9 @@ droidrun run "Open YouTube"
```python
# In scripts
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(goal="Open YouTube", config=config)
```
@@ -309,9 +309,9 @@ droidrun run "Find contacts from California and export to CSV" --reasoning
```python
# In scripts - Option 1: Override config
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig, AgentConfig
from droidrun.config_manager.config_manager import DroidrunConfig, AgentConfig
config = DroidRunConfig()
config = DroidrunConfig()
config.agent.reasoning = True
agent = DroidAgent(goal="...", config=config)
@@ -323,15 +323,14 @@ agent = DroidAgent(goal="...", config=config, agent_config=agent_config)
## Next Steps
Now that you've got DroidRun v4 running, explore these topics:
Now that you've got Droidrun v4 running, explore these topics:
### Core Concepts
- [Architecture](/v4/concepts/architecture) - Multi-agent system architecture
- [Workflow Architecture](/v4/concepts/workflow-architecture) - Event-driven coordination
- [Event Streaming](/v4/concepts/event-streaming) - Real-time execution monitoring
- [Events and Workflows](/v4/concepts/events-and-workflows) - Event-driven coordination and real-time execution monitoring
### Configuration
- [Configuration System](/v4/guides/configuration) - Complete configuration guide
- [Configuration System](/v4/sdk/configuration) - Complete configuration guide
- [Custom Tools & Credentials](/v4/guides/custom-tools-credentials) - Extend functionality
- [Custom Variables](/v4/guides/custom-variables) - Pass custom data to agents
- [Structured Output](/v4/guides/structured-output) - Extract structured data
@@ -342,10 +341,10 @@ Now that you've got DroidRun v4 running, explore these topics:
### Advanced Topics
- [Vision Mode](/v4/concepts/architecture#vision-configuration) - Screenshot processing
- [Safe Execution](/v4/guides/configuration#safe-execution) - Secure code execution
- [Tracing & Telemetry](/v4/guides/configuration#tracing-settings) - Debugging and monitoring
- [Custom Prompts](/v4/guides/configuration#prompt-customization) - Customize agent behavior
- [Safe Execution](/v4/sdk/configuration#safe-execution) - Secure code execution
- [Tracing & Telemetry](/v4/sdk/configuration#tracing-settings) - Debugging and monitoring
- [Custom Prompts](/v4/sdk/configuration#prompt-customization) - Customize agent behavior
---
**Welcome to DroidRun v4!** The config-driven architecture gives you complete control over agent behavior, making it easier than ever to build powerful device automation workflows.
**Welcome to Droidrun v4!** The config-driven architecture gives you complete control over agent behavior, making it easier than ever to build powerful device automation workflows.
+49
View File
@@ -0,0 +1,49 @@
---
title: 'SDK Reference'
description: 'Complete API reference for Droidrun v4 components and tools'
---
## Overview
The Droidrun SDK provides a comprehensive set of APIs for building mobile automation workflows with AI agents. This reference documentation covers all major components and tools.
---
## Core Components
<CardGroup cols={2}>
<Card title="DroidAgent" icon="robot" href="/v4/sdk/droid-agent">
Main agent coordinator with multi-agent orchestration
</Card>
<Card title="ADB Tools" icon="mobile" href="/v4/sdk/adb-tools">
Android device control via ADB
</Card>
<Card title="iOS Tools" icon="apple" href="/v4/sdk/ios-tools">
iOS device control and automation
</Card>
<Card title="Base Tools" icon="wrench" href="/v4/sdk/base-tools">
Abstract base classes for tool implementations
</Card>
</CardGroup>
---
## Configuration
<Card title="Configuration API" icon="gear" href="/v4/sdk/configuration">
DroidrunConfig API and YAML configuration reference
</Card>
---
## API Documentation
Detailed API documentation for each component is available in the sections linked above. Each page includes:
- Class/function signatures
- Parameter descriptions
- Return types
- Usage examples
- Best practices
For conceptual guides and tutorials, see the [Guides](/v4/guides/overview) section.
+7 -7
View File
@@ -14,7 +14,7 @@ class AdbTools(Tools)
Core UI interaction tools for Android device control.
AdbTools provides a comprehensive interface for interacting with Android devices through ADB (Android Debug Bridge). It supports both TCP communication and content provider modes for device communication via the DroidRun Portal app.
AdbTools provides a comprehensive interface for interacting with Android devices through ADB (Android Debug Bridge). It supports both TCP communication and content provider modes for device communication via the Droidrun Portal app.
<a id="droidrun.tools.adb.AdbTools.__init__"></a>
@@ -67,7 +67,7 @@ tools = AdbTools(
```
**Notes:**
- Automatically sets up the DroidRun Portal keyboard on initialization via `setup_keyboard()`
- Automatically sets up the Droidrun Portal keyboard on initialization via `setup_keyboard()`
- Creates a PortalClient instance that handles TCP/content provider communication
- Device serial can be emulator name, USB serial, or TCP/IP address:port
@@ -293,7 +293,7 @@ result = tools.input_text("Hello\nWorld") # Multiline text
**Notes:**
- Always ensure a text field is focused before inputting text (use `tap_by_index()` or set `index` parameter)
- Uses the DroidRun Portal app keyboard for reliable text input via PortalClient
- Uses the Droidrun Portal app keyboard for reliable text input via PortalClient
- Supports Unicode characters and special characters including non-ASCII
- If `index != -1`, automatically taps the element first before inputting text
- Call `get_state()` first to populate element cache if using `index` parameter
@@ -803,7 +803,7 @@ tools.complete(success=False, reason="Could not find contact 'John' in contacts
## Notes
- **Portal app required**: The DroidRun Portal app must be installed and accessibility service enabled on the device
- **Portal app required**: The Droidrun Portal app must be installed and accessibility service enabled on the device
- **TCP vs Content Provider**: TCP is faster but requires port forwarding (`adb forward tcp:8080 tcp:8080`). Content provider is the fallback mode using ADB shell commands.
- **Element caching**: Always call `get_state()` before using `tap_by_index()` or `tap()` to populate the element cache
- **Trajectory recording**: When `save_trajectories="action"`, screenshots and UI states are automatically captured for each UI action via the `@Tools.ui_action` decorator
@@ -843,7 +843,7 @@ for element in state['a11y_tree']:
break
# Input search query
result = tools.input_text("DroidRun framework")
result = tools.input_text("Droidrun framework")
print(result)
# Press enter key
@@ -857,10 +857,10 @@ with open("search_result.png", "wb") as f:
f.write(screenshot)
# Remember result for future context
tools.remember("Searched for DroidRun framework in Chrome")
tools.remember("Searched for Droidrun framework in Chrome")
# Complete task
tools.complete(success=True, reason="Successfully searched for DroidRun in Chrome")
tools.complete(success=True, reason="Successfully searched for Droidrun in Chrome")
# Check completion status
print(f"Task finished: {tools.finished}")
+6 -6
View File
@@ -350,13 +350,13 @@ Tools instances are passed to agents and provide the atomic actions for device c
```python
from droidrun import DroidAgent
from droidrun.tools import AdbTools
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Create tools instance
tools = AdbTools(serial="emulator-5554")
# Create config
config = DroidRunConfig()
config = DroidrunConfig()
# Pass to agent
agent = DroidAgent(
@@ -544,7 +544,7 @@ tools.swipe(100, 500, 100, 100) # Logs + captures screenshot
## See Also
- [AdbTools API](/docs/v4/sdk/adb-tools) - Android implementation
- [IOSTools API](/docs/v4/sdk/ios-tools) - iOS implementation
- [DroidAgent API](/docs/v4/sdk/droid-agent) - Agent integration
- [Custom Tools Guide](/docs/v4/guides/custom-tools-credentials) - Creating custom tools
- [AdbTools API](/v4/sdk/adb-tools) - Android implementation
- [IOSTools API](/v4/sdk/ios-tools) - iOS implementation
- [DroidAgent API](/v4/sdk/droid-agent) - Agent integration
- [Custom Tools Guide](/v4/guides/custom-tools-credentials) - Creating custom tools
+693
View File
@@ -0,0 +1,693 @@
---
title: 'Configuration Reference'
description: 'Complete DroidAgent configuration guide - all parameters, minimal examples'
---
## Quick Start
```python
from droidrun import DroidAgent
# Minimal (uses defaults)
agent = DroidAgent(goal="Open settings")
result = await agent.run()
```
---
## DroidAgent Parameters
### Required
```python
DroidAgent(
goal="Your task", # REQUIRED: Task description
)
```
### Optional Parameters
| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `config` | `DroidrunConfig \| None` | `None` | Full config object (loads LLMs from profiles if `llms` not provided) |
| `llms` | `dict[str, LLM] \| LLM \| None` | `None` | LLM(s) - dict for per-agent, single LLM for all, or None to load from config |
| `agent_config` | `AgentConfig \| None` | `None` | Agent behavior settings (overrides config.agent) |
| `device_config` | `DeviceConfig \| None` | `None` | Device connection settings (overrides config.device) |
| `tools` | `Tools \| ToolsConfig \| None` | `None` | Tools instance or config (overrides config.tools) |
| `logging_config` | `LoggingConfig \| None` | `None` | Logging settings (overrides config.logging) |
| `tracing_config` | `TracingConfig \| None` | `None` | Tracing settings (overrides config.tracing) |
| `telemetry_config` | `TelemetryConfig \| None` | `None` | Telemetry settings (overrides config.telemetry) |
| `custom_tools` | `dict \| None` | `None` | Custom tool definitions |
| `credentials` | `CredentialsConfig \| dict \| None` | `None` | Credentials config or dict of secrets |
| `variables` | `dict \| None` | `None` | Custom variables accessible during execution |
| `output_model` | `Type[BaseModel] \| None` | `None` | Pydantic model for structured output extraction |
| `prompts` | `dict[str, str] \| None` | `None` | Custom Jinja2 prompt templates (NOT file paths) |
| `timeout` | `int` | `1000` | Workflow timeout in seconds |
---
## Configuration Classes
### AgentConfig
```python
from droidrun.config_manager.config_manager import AgentConfig, CodeActConfig, ManagerConfig, ExecutorConfig, ScripterConfig, AppCardConfig
AgentConfig(
# Core settings
max_steps=15, # Max execution steps
reasoning=False, # Enable Manager/Executor workflow
after_sleep_action=1.0, # Wait after actions (seconds)
wait_for_stable_ui=0.3, # Wait for UI to stabilize (seconds)
prompts_dir="config/prompts", # Prompt templates directory
# Sub-configs
codeact=CodeActConfig(...),
manager=ManagerConfig(...),
executor=ExecutorConfig(...),
scripter=ScripterConfig(...),
app_cards=AppCardConfig(...),
)
```
**CodeActConfig**
```python
CodeActConfig(
vision=False, # Enable screenshots
system_prompt="system.jinja2", # Filename in prompts_dir/codeact/
user_prompt="user.jinja2", # Filename in prompts_dir/codeact/
safe_execution=False, # Restrict imports/builtins
)
```
**ManagerConfig**
```python
ManagerConfig(
vision=False, # Enable screenshots
system_prompt="system.jinja2", # Filename in prompts_dir/manager/
)
```
**ExecutorConfig**
```python
ExecutorConfig(
vision=False, # Enable screenshots
system_prompt="system.jinja2", # Filename in prompts_dir/executor/
)
```
**ScripterConfig**
```python
ScripterConfig(
enabled=True, # Enable off-device Python execution
max_steps=10, # Max scripter steps
execution_timeout=30.0, # Code block timeout (seconds)
system_prompt_path="system.jinja2", # Filename in prompts_dir/scripter/
safe_execution=False, # Restrict imports/builtins
)
```
**AppCardConfig**
```python
AppCardConfig(
enabled=True, # Enable app-specific instructions
mode="local", # "local" | "server" | "composite"
app_cards_dir="config/app_cards", # Directory for app card files
server_url=None, # Server URL (for server/composite modes)
server_timeout=2.0, # Server request timeout (seconds)
server_max_retries=2, # Server retry attempts
)
```
---
### DeviceConfig
```python
from droidrun.config_manager.config_manager import DeviceConfig
DeviceConfig(
serial=None, # Device serial/IP (None = auto-detect)
platform="android", # "android" or "ios"
use_tcp=False, # TCP vs content provider communication
)
```
---
### LoggingConfig
```python
from droidrun.config_manager.config_manager import LoggingConfig
LoggingConfig(
debug=False, # Enable debug logs
save_trajectory="none", # "none" | "step" | "action"
rich_text=False, # Rich text formatting in logs
)
```
---
### TracingConfig
```python
from droidrun.config_manager.config_manager import TracingConfig
TracingConfig(
enabled=False, # Enable Arize Phoenix tracing
)
```
---
### TelemetryConfig
```python
from droidrun.config_manager.config_manager import TelemetryConfig
TelemetryConfig(
enabled=True, # Enable anonymous telemetry
)
```
---
### ToolsConfig
```python
from droidrun.config_manager.config_manager import ToolsConfig
ToolsConfig(
allow_drag=False, # Enable drag tool
)
```
---
### CredentialsConfig
```python
from droidrun.config_manager.config_manager import CredentialsConfig
CredentialsConfig(
enabled=False, # Enable credential manager
file_path="credentials.yaml", # Path to credentials file
)
```
---
### SafeExecutionConfig
```python
from droidrun.config_manager.safe_execution import SafeExecutionConfig
SafeExecutionConfig(
# Imports
allow_all_imports=False, # Allow all imports (ignores allowed_modules)
allowed_modules=[], # Allowed module names (e.g., ["json", "requests"])
blocked_modules=[ # Blocked modules (takes precedence)
"os", "sys", "subprocess", "shutil", "pathlib", "pty", "fcntl",
"resource", "pickle", "shelve", "marshal", "imp", "importlib",
"ctypes", "code", "codeop", "tempfile", "glob", "socket",
"socketserver", "asyncio"
],
# Builtins
allow_all_builtins=False, # Allow all builtins (ignores allowed_builtins)
allowed_builtins=[], # Allowed builtin names (empty = safe defaults)
blocked_builtins=[ # Blocked builtins (takes precedence)
"open", "compile", "exec", "eval", "__import__",
"breakpoint", "exit", "quit", "input"
],
)
```
---
## LLM Configuration
### Single LLM (All Agents)
```python
from llama_index.llms.gemini import Gemini
llm = Gemini(model="models/gemini-2.5-pro", temperature=0.2)
agent = DroidAgent(goal="...", llms=llm)
```
### Per-Agent LLMs
```python
from llama_index.llms.openai import OpenAI
from llama_index.llms.gemini import Gemini
agent = DroidAgent(
goal="...",
llms={
"manager": OpenAI(model="gpt-4o"), # Planning
"executor": Gemini(model="models/gemini-2.5-flash"), # Action selection
"codeact": Gemini(model="models/gemini-2.5-pro"), # Code generation
"text_manipulator": Gemini(model="models/gemini-2.5-flash"), # Text input
"app_opener": OpenAI(model="gpt-4o-mini"), # App launching
"scripter": Gemini(model="models/gemini-2.5-flash"), # Off-device scripts
"structured_output": Gemini(model="models/gemini-2.5-flash"), # Output extraction
}
)
```
**LLM Keys:**
- `manager` - Planning (reasoning mode only)
- `executor` - Action selection (reasoning mode only)
- `codeact` - Code generation (direct mode)
- `scripter` - Off-device Python execution
- `text_manipulator` - Text input helper
- `app_opener` - App launching helper
- `structured_output` - Final output extraction
---
## Custom Tools
```python
def my_tool(param: str) -> str:
"""Tool description."""
return f"Result: {param}"
agent = DroidAgent(
goal="...",
custom_tools={
"my_tool": {
"signature": "my_tool(param: str) -> str",
"description": "Tool description",
"function": my_tool
}
}
)
```
---
## Credentials
### Dict Format (Recommended)
```python
agent = DroidAgent(
goal="...",
credentials={
"USERNAME": "alice@example.com",
"PASSWORD": "secret123"
}
)
# Agent can call get_username() and get_password()
```
### Config Format
```python
from droidrun.config_manager.config_manager import CredentialsConfig
agent = DroidAgent(
goal="...",
credentials=CredentialsConfig(
enabled=True,
file_path="config/credentials.yaml"
)
)
```
---
## Custom Variables
```python
agent = DroidAgent(
goal="...",
variables={
"api_url": "https://api.example.com",
"user_id": "12345",
"custom_data": {"key": "value"}
}
)
# Access in shared_state.custom_variables
```
---
## Structured Output
```python
from pydantic import BaseModel
class FlightInfo(BaseModel):
airline: str
flight_number: str
confirmation_code: str
agent = DroidAgent(
goal="Book a flight and extract details",
output_model=FlightInfo
)
result = await agent.run()
print(result.structured_output.airline) # Typed output
```
---
## Custom Prompts
```python
custom_prompt = """
You are an expert mobile agent.
Goal: {{ instruction }}
Be precise and efficient.
"""
agent = DroidAgent(
goal="...",
prompts={
"codeact_system": custom_prompt,
"codeact_user": "...",
"manager_system": "...",
"executor_system": "...",
"scripter_system": "..."
}
)
```
**Template Variables:**
- `{{ instruction }}` - User's goal
- `{{ device_date }}` - Device date/time
- `{{ app_card }}` - App-specific instructions
- `{{ state }}` - Device state
- `{{ history }}` - Action history
---
## Complete Example
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import (
AgentConfig, CodeActConfig, DeviceConfig, LoggingConfig, TracingConfig
)
from llama_index.llms.openai import OpenAI
from llama_index.llms.gemini import Gemini
from pydantic import BaseModel
# Structured output
class Output(BaseModel):
name: str
value: int
# Custom tool
def send_email(to: str, subject: str) -> str:
"""Send email."""
return f"Sent to {to}"
agent = DroidAgent(
goal="Complex task",
# LLMs
llms={
"manager": OpenAI(model="gpt-4o"),
"executor": Gemini(model="models/gemini-2.5-flash"),
"codeact": Gemini(model="models/gemini-2.5-pro")
},
# Agent behavior
agent_config=AgentConfig(
max_steps=30,
reasoning=True,
after_sleep_action=1.5,
codeact=CodeActConfig(vision=True, safe_execution=True)
),
# Device
device_config=DeviceConfig(
serial="emulator-5554",
platform="android",
use_tcp=False
),
# Logging
logging_config=LoggingConfig(
debug=True,
save_trajectory="action"
),
# Tracing
tracing_config=TracingConfig(enabled=True),
# Custom tools
custom_tools={
"send_email": {
"signature": "send_email(to: str, subject: str) -> str",
"description": "Send email",
"function": send_email
}
},
# Credentials
credentials={"USERNAME": "alice", "PASSWORD": "secret"},
# Variables
variables={"api_url": "https://api.example.com"},
# Structured output
output_model=Output,
# Timeout
timeout=600
)
result = await agent.run()
```
---
## YAML Config (CLI)
For CLI usage, create `config.yaml`:
```yaml
agent:
max_steps: 15
reasoning: false
after_sleep_action: 1.0
wait_for_stable_ui: 0.3
prompts_dir: config/prompts
codeact:
vision: false
system_prompt: system.jinja2
user_prompt: user.jinja2
safe_execution: false
manager:
vision: false
system_prompt: system.jinja2
executor:
vision: false
system_prompt: system.jinja2
scripter:
enabled: true
max_steps: 10
execution_timeout: 30.0
system_prompt_path: system.jinja2
safe_execution: false
app_cards:
enabled: true
mode: local
app_cards_dir: config/app_cards
server_url: null
server_timeout: 2.0
server_max_retries: 2
llm_profiles:
manager:
provider: GoogleGenAI
model: models/gemini-2.5-pro
temperature: 0.2
kwargs:
max_tokens: 8192
executor:
provider: GoogleGenAI
model: models/gemini-2.5-flash
temperature: 0.1
kwargs:
max_tokens: 4096
codeact:
provider: GoogleGenAI
model: models/gemini-2.5-pro
temperature: 0.2
kwargs:
max_tokens: 8192
text_manipulator:
provider: GoogleGenAI
model: models/gemini-2.5-flash
temperature: 0.3
app_opener:
provider: OpenAI
model: gpt-4o-mini
temperature: 0.0
scripter:
provider: GoogleGenAI
model: models/gemini-2.5-flash
temperature: 0.1
structured_output:
provider: GoogleGenAI
model: models/gemini-2.5-flash
temperature: 0.0
device:
serial: null
platform: android
use_tcp: false
telemetry:
enabled: true
tracing:
enabled: false
logging:
debug: false
save_trajectory: none
rich_text: false
safe_execution:
allow_all_imports: false
allowed_modules: []
blocked_modules:
- os
- sys
- subprocess
- shutil
- pathlib
- pty
- fcntl
- resource
- pickle
- shelve
- marshal
- imp
- importlib
- ctypes
- code
- codeop
- tempfile
- glob
- socket
- socketserver
- asyncio
allow_all_builtins: false
allowed_builtins: []
blocked_builtins:
- open
- compile
- exec
- eval
- __import__
- breakpoint
- exit
- quit
- input
tools:
allow_drag: false
credentials:
enabled: false
file_path: credentials.yaml
```
---
## CLI Overrides
```bash
# Override agent settings
droidrun run "Task" --steps 30 --reasoning --vision
# Override device
droidrun run "Task" --device emulator-5554 --tcp
# Override LLM (applies to ALL agents)
droidrun run "Task" --provider GoogleGenAI --model models/gemini-2.5-flash
# Override logging
droidrun run "Task" --debug --save-trajectory action --tracing
# Custom config file
droidrun run "Task" --config /path/to/config.yaml
```
**All CLI Flags:**
- `--config PATH` - Custom config file
- `--device SERIAL` - Device serial/IP
- `--provider PROVIDER` - LLM provider (OpenAI, Ollama, Anthropic, GoogleGenAI, DeepSeek)
- `--model MODEL` - LLM model name
- `--temperature FLOAT` - LLM temperature
- `--steps INT` - Max steps
- `--base_url URL` - API base URL (for Ollama/OpenRouter)
- `--api_base URL` - API base URL (for OpenAI-like)
- `--vision/--no-vision` - Enable/disable vision for all agents
- `--reasoning/--no-reasoning` - Enable/disable reasoning mode
- `--tracing/--no-tracing` - Enable/disable tracing
- `--debug/--no-debug` - Enable/disable debug logs
- `--tcp/--no-tcp` - Enable/disable TCP communication
- `--save-trajectory none|step|action` - Trajectory saving level
- `--ios` - Run on iOS device
---
## Configuration Priority
When multiple sources are provided, DroidAgent uses this priority:
1. **Direct parameters** (highest priority)
```python
agent = DroidAgent(goal="...", agent_config=AgentConfig(max_steps=30))
```
2. **Config object**
```python
agent = DroidAgent(goal="...", config=my_config)
```
3. **Defaults** (lowest priority)
**Example:**
```python
config = DroidrunConfig(agent=AgentConfig(max_steps=20))
agent = DroidAgent(
goal="...",
config=config, # max_steps=20 from config
agent_config=AgentConfig(max_steps=30) # Overrides to 30
)
```
---
## Environment Variables
Set API keys via environment variables:
```bash
export GOOGLE_API_KEY=your-key
export OPENAI_API_KEY=your-key
export ANTHROPIC_API_KEY=your-key
export DEEPSEEK_API_KEY=your-key
export DROIDRUN_CONFIG=/path/to/config.yaml # Custom config path
```
+32 -33
View File
@@ -25,7 +25,7 @@ A wrapper class that coordinates between agents to achieve a user's goal.
```python
def __init__(
goal: str,
config: DroidRunConfig | None = None,
config: DroidrunConfig | None = None,
llms: dict[str, LLM] | LLM | None = None,
agent_config: AgentConfig | None = None,
device_config: DeviceConfig | None = None,
@@ -47,7 +47,7 @@ Initialize the DroidAgent wrapper.
**Arguments**:
- `goal` _str_ - User's goal or command to execute
- `config` _DroidRunConfig | None_ - Full configuration object (required if llms not provided). Contains agent settings, LLM profiles, device config, and more. If provided, individual config overrides (agent_config, device_config, etc.) take precedence.
- `config` _DroidrunConfig | None_ - Full configuration object (required if llms not provided). Contains agent settings, LLM profiles, device config, and more. If provided, individual config overrides (agent_config, device_config, etc.) take precedence.
- `llms` _dict[str, LLM] | LLM | None_ - Optional LLM configuration:
- `dict[str, LLM]`: Agent-specific LLMs with keys: "manager", "executor", "codeact", "text_manipulator", "app_opener", "scripter", "structured_output"
- `LLM`: Single LLM instance used for all agents
@@ -75,14 +75,14 @@ Initialize the DroidAgent wrapper.
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize with default config
config = DroidRunConfig()
config = DroidrunConfig()
# Create agent (LLMs loaded from config.llm_profiles)
agent = DroidAgent(
goal="Open Chrome and search for DroidRun",
goal="Open Chrome and search for Droidrun",
config=config
)
@@ -94,14 +94,14 @@ result = await agent.run()
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Load config from config.yaml
config = DroidRunConfig.from_yaml("config.yaml")
config = DroidrunConfig.from_yaml("config.yaml")
# Create agent (LLMs loaded from config.llm_profiles)
agent = DroidAgent(
goal="Open Chrome and search for DroidRun",
goal="Open Chrome and search for Droidrun",
config=config
)
@@ -113,17 +113,17 @@ result = await agent.run()
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
from llama_index.llms.openai import OpenAI
from llama_index.llms.anthropic import Anthropic
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Create custom LLMs
llms = {
"manager": Anthropic(model="claude-sonnet-4-5", temperature=0.2),
"executor": Anthropic(model="claude-sonnet-4-5", temperature=0.1),
"manager": Anthropic(model="claude-sonnet-4-5-latest", temperature=0.2),
"executor": Anthropic(model="claude-sonnet-4-5-latest", temperature=0.1),
"codeact": OpenAI(model="gpt-4o", temperature=0.2),
"text_manipulator": OpenAI(model="gpt-4o-mini", temperature=0.3),
"app_opener": OpenAI(model="gpt-4o-mini", temperature=0.0),
@@ -145,11 +145,11 @@ result = await agent.run()
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
from llama_index.llms.openai import OpenAI
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Use same LLM for all agents
llm = OpenAI(model="gpt-4o", temperature=0.2)
@@ -167,10 +167,10 @@ result = await agent.run()
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Define custom tool
def search_database(query: str) -> str:
@@ -206,11 +206,11 @@ result = await agent.run()
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
from pydantic import BaseModel, Field
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Define output schema
class WeatherInfo(BaseModel):
@@ -256,10 +256,10 @@ Run the DroidAgent workflow.
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Create and run agent
agent = DroidAgent(goal="...", config=config)
@@ -274,10 +274,10 @@ print(f"Steps: {result.steps}")
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(goal="...", config=config)
@@ -328,7 +328,7 @@ DroidAgent emits various events during execution:
## Configuration
DroidAgent uses a hierarchical configuration system. See the [Configuration Guide](/docs/v4/guides/configuration) for details.
DroidAgent uses a hierarchical configuration system. See the [Configuration Guide](/v4/sdk/configuration) for details.
**Key configuration options:**
@@ -365,19 +365,18 @@ tracing:
**Custom Tools instance:**
```python
from droidrun import DroidAgent, AdbTools
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun import DroidAgent, DeviceConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Pre-configure tools
tools = AdbTools(serial="emulator-5554", use_tcp=True)
device_config = DeviceConfig(serial="emulator-5554", use_tcp=True)
agent = DroidAgent(
goal="Open settings",
config=config,
tools=tools # Use pre-configured tools
device_config=device_config
)
result = await agent.run()
@@ -387,10 +386,10 @@ result = await agent.run()
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
agent = DroidAgent(
goal="Complete task using context",
@@ -411,10 +410,10 @@ Variables are accessible in shared_state.custom_variables throughout execution a
```python
from droidrun import DroidAgent
from droidrun.config_manager.config_manager import DroidRunConfig
from droidrun.config_manager.config_manager import DroidrunConfig
# Initialize config
config = DroidRunConfig()
config = DroidrunConfig()
# Override default prompts with custom Jinja2 templates
custom_prompts = {
+5 -5
View File
@@ -12,7 +12,7 @@ title: IOSTools
class IOSTools(Tools)
```
Core UI interaction tools for iOS device control via the DroidRun iOS Portal app.
Core UI interaction tools for iOS device control via the Droidrun iOS Portal app.
**Status**: iOS support is in beta with limited functionality compared to AdbTools.
@@ -59,7 +59,7 @@ tools = IOSTools(
**Setup Requirements:**
1. Install DroidRun iOS Portal app on device
1. Install Droidrun iOS Portal app on device
2. Launch Portal app (starts HTTP server)
3. Connect device and computer to same network
4. Use displayed URL to initialize IOSTools
@@ -612,7 +612,7 @@ for elem in state['a11y_tree']:
if elem['type'] == 'TextField' and 'message' in elem['label'].lower():
tools.tap_by_index(elem['index'])
break
tools.input_text("Hello from DroidRun!")
tools.input_text("Hello from Droidrun!")
# Send
state = tools.get_state()
@@ -630,5 +630,5 @@ tools.complete(True, "Successfully sent iMessage")
## See Also
- [AdbTools](/docs/v4/sdk/adb-tools) - Android device control with full functionality
- [Tools Base Class](/docs/v4/sdk/base-tools) - Abstract base class reference
- [AdbTools](/v4/sdk/adb-tools) - Android device control with full functionality
- [Tools Base Class](/v4/sdk/base-tools) - Abstract base class reference
+1 -1
View File
@@ -6,7 +6,7 @@ pydoc-markdown
# rename .md to .mdx
# Rename all .md files to .mdx in the docs/v3/sdk directory
find docs/v3/sdk -name "*.md" -exec sh -c 'mv "$1" "${1%.md}.mdx"' _ {} \;
find docs/v4/sdk -name "*.md" -exec sh -c 'mv "$1" "${1%.md}.mdx"' _ {} \;
# update docs/v3/.generated-files.txt with the new files extension
# Set sed in-place flag for macOS and Linux compatibility
+1 -1
View File
@@ -120,7 +120,7 @@ type = "crossref"
[tool.pydoc-markdown.renderer]
type = "hugo"
build_directory = "docs"
content_directory = "v3/sdk"
content_directory = "v4/sdk"
clean_render = true