Getting Started

PNP Compute is an open-source ecosystem that turns a dedicated machine into an AI-powered workstation you can control from your phone. This guide will get you up and running in under 10 minutes.

Installation

Clone the repository and install the Python dependencies:

# GitHub repo — coming soon
cd neurosdk
python3 -m venv venv && source venv/bin/activate
pip install -r requirements.txt

You will also need a local LiveKit server for WebRTC signaling and streaming:

# Install LiveKit (Linux)
curl -sSL https://get.livekit.io | bash

# Copy the example config
cp livekit.yaml.example livekit.yaml

Configuration

Create a .env file in the project root with your API keys:

OPENAI_API_KEY=sk-...
ELEVENLABS_API_KEY=...
SARVAM_API_KEY=...
LIVEKIT_URL=ws://localhost:7880
LIVEKIT_API_KEY=devkey
LIVEKIT_API_SECRET=secret

Edit livekit.yaml to set your machine's public IP address under the rtc.use_external_ip field. This is required for the mobile app to connect over the network.

Running the Server

# Install vendored NeuroLang library
pip install -e ./neurolang

# Start LiveKit in the background (optional)
livekit-server --config livekit.yaml &

# Start the PNP Compute server
python neurosdk/server.py
# Server is now running at http://127.0.0.1:7000

# Start the IDE backend (optional, port 8000)
python3 neurosdk/scripts/ide_server.py

# Start the web UI
cd neuro_web && npm install && npm run dev
# Open http://localhost:3000

To build and install the Android app:

cd neuro_mobile
./gradlew assembleDebug
# Install the APK on your device
# Set server URL in Settings → http://<your-machine-ip>:7000

Mobile Application

The Android application is built with Jetpack Compose and serves as the unified interface for the entire PNP Compute ecosystem. It handles chat, voice, agent switching, and full remote desktop control.

UI Elements

  • Agent Selector — Tap the agent pill in the header to switch between Neuro, NL Dev, OpenClaw, OpenCode, and NeuroUpwork. Each agent has its own conversation context.
  • Chat Interface — Full conversation history with text messages, voice waveform indicators, and per-message TTS playback buttons.
  • Voice Typing — Hold the microphone button to dictate. The transcription populates the input field in real time.
  • Tab Bar — Supports multiple conversations per agent, each tracked independently with its own chat history.
  • Side Drawer — Access remote desktop mode, settings, and overlay controls from the hamburger menu.

Remote Desktop Mode

Tap "Remote Desktop" in the side drawer to enter fullscreen landscape mode. The desktop screen is streamed in real time via WebRTC.

  • Right Sidebar — Microphone toggle, mouse mode toggle, monitor switch (multi-display), and exit fullscreen button.
  • Left Toolbar — A draggable floating panel with voice typing, keyboard toggle, scroll/click/focus mode switches, and agent switcher.
  • Zoom & Pan — Pinch to zoom up to 10× on any area of the screen. Drag to pan when zoomed.

Touch Gestures

GestureAction
Single-finger dragMove mouse cursor (relative)
TapLeft click
Double tapDouble click
Two-finger dragScroll (vertical/horizontal)
Two-finger tapRight click
PinchZoom in/out on screen

Voice Pipeline

The voice system is optimized for low latency across the full round trip:

Microphone
    ↓
Silero VAD (Voice Activity Detection)
    ↓
Sarvam Streaming STT (Speech-to-Text)
    ↓
Brain (LLM reasoning + skill execution)
    ↓
ElevenLabs TTS (Text-to-Speech)
    ↓
Speaker

All audio travels through LiveKit, the same transport used for the desktop video stream. This keeps latency low and avoids separate connection overhead.

Neuro Framework — The Brain

The Brain is the central intelligence layer. It implements a Reason + Act (ReAct) loop that processes every incoming message — whether text or voice — and decides how to respond.

Smart Router

The first thing the Brain does is classify the request. In a single LLM call, the Smart Router decides:

  1. Direct reply — The request is conversational (e.g., "What time is it?"). Respond immediately without invoking any skills.
  2. Skill execution — The request requires action (e.g., "Lock my screen"). Identify the correct skill and pass it to the Planner.

This classification keeps simple conversations fast (no planning overhead) and complex tasks structured.

Planner & DAG Execution

For multi-step requests, the Planner breaks the goal into a Directed Acyclic Graph (DAG) of skill calls. Each node in the graph represents one skill invocation. Outputs from earlier nodes feed into later ones.

User: "Find my latest screenshot and describe what's on screen"

DAG:
  Node 1: list_files(dir="~/Screenshots", sort="newest")
  Node 2: read_file(path=Node1.output[0])
  Node 3: describe_image(image=Node2.output)
  Node 4: reply(text=Node3.output)

The Executor runs this graph node by node, handling errors and retries along the way.

Skills System

Skills are the atomic units of capability in PNP Compute. Each skill is a self-contained folder:

skills/lock_screen/
├── conf.json    # Name, description, parameter schema
└── code.py      # execute() function

The Brain discovers skills at runtime by scanning the skills directory. To add a new skill, simply drop a new folder in — no restart required.

Built-in skill categories:

CategoryExamples
DesktopLock screen, take screenshot, open file explorer, move mouse
CodeRead/write/diff files, scan projects, generate patches
UpworkList jobs, analyze fit, draft proposals, capture screenshots
BrowserDelegate tasks to OpenClaw for web interaction
MetaList skills, create new skills, edit existing skills at runtime

Profiles

Profiles control which skills and prompt are active at any given time. Each agent owns a default profile.

ProfileUse Case
generalDefault. All general-purpose skills available.
code_devCode editing and project management focus.
neuro_devFull meta-programming — build and edit neuros themselves.
neurolang_devAuthoring NeuroLang flows; paired with the nl_dev agent.

Switch via API: POST /api/profile/switch { "profile": "neurolang_dev" }. Inspect with /api/profile/active and /api/profile/list.

Multi-Agent

Three primitives compose every multi-agent pattern in PNP Compute.

Meeting Rooms

An isolated real-time collaboration space. Multiple agents share a transcript with a mediator picking the next speaker round-robin. Backed by core/rooms.py, core/rooms_db.py, and four neuros: room_create, room_post, room_close, room_mediator. UI: neuro_web/components/rooms/RoomPanel.tsx. Endpoint: /api/rooms.

agent.talk

Direct typed messages between agents. Depth-guarded (MAX=4) via TalkDepthExceeded. Two neuros: agent_talk, agent_list. The substrate beneath rooms.

Schedules

Cron-style triggers persisted to schedules.db via APScheduler. Three neuros: schedule_run, schedule_cancel, schedule_list. Endpoint: /api/schedules.

NeuroLang Integration

NeuroLang is vendored at ./neurolang/ (Phase 1.9, 172 tests passing). Install with pip install -e ./neurolang.

The nl_dev agent wraps NeuroLang's compile_source / propose_plan / decompile_summary through seven nl_* neuros: nl_planner, nl_propose, nl_compile, nl_save, nl_run, nl_summary, nl_reply. Compiled flows land in ~/.neurolang/neuros/.

See the NeuroLang page for the language itself.

API Reference

Chat

MethodEndpointDescription
POST/chatSend a message to the active agent
POST/conversationCreate a new conversation
GET/conversations?agent_id=List conversations by agent

Agents

MethodEndpointDescription
GET/agentsList all running agents
POST/agents/{type}Switch to a specific agent type

Streaming

MethodEndpointDescription
POST/stream/startBegin desktop screen streaming
GET/voice/tokenGet LiveKit token for voice sessions

Input Control

MethodEndpointDescription
POST/mouse/moveMove mouse cursor (relative)
POST/mouse/clickClick at current position
POST/mouse/scrollScroll vertically/horizontally
POST/keyboard/sendSend keystrokes and key combos

Full endpoint details are available in GitHub · coming soon. Also see the API summary on the architecture page.