feat(ai): native host-side tool-calling harness (Phase 2)
CI / rust client (hh) (macos-latest) (push) Has been cancelled
CI / rust client (hh) (ubuntu-latest) (push) Has been cancelled
CI / rust coverage (push) Has been cancelled
CI / python server (3.10) (push) Has been cancelled
CI / python server (3.11) (push) Has been cancelled
CI / python server (3.12) (push) Has been cancelled
CI / headless e2e smoke (push) Has been cancelled
CI / dependency audit (push) Has been cancelled
CI / secret scanning (push) Has been cancelled
CI / rust client (hh) (macos-latest) (push) Has been cancelled
CI / rust client (hh) (ubuntu-latest) (push) Has been cancelled
CI / rust coverage (push) Has been cancelled
CI / python server (3.10) (push) Has been cancelled
CI / python server (3.11) (push) Has been cancelled
CI / python server (3.12) (push) Has been cancelled
CI / headless e2e smoke (push) Has been cancelled
CI / dependency audit (push) Has been cancelled
CI / secret scanning (push) Has been cancelled
Implement the bounded native harness from docs/spec-native-harness.md §1.3 and
make it the default granted-!task path. The model runs host-side (no container→
host Ollama hop); only its tool calls exec in the sandbox.
providers.py:
- OllamaProvider.complete_with_tools(system, messages, tools) -> (text, calls):
one non-streaming /api/chat turn with a `tools` schema; parses message.tool_calls
(dict or JSON-string arguments). Caches tool capability (_tools_ok / supports_tools).
- ToolsUnsupported raised when the model rejects `tools` ("does not support tools").
bridge.py:
- NATIVE_SYSTEM + a 3-tool schema (run_shell / write_file / read_file), turn/byte caps.
- _run_native: seed transcript window + task → loop up to max_turns; exec each tool
call in the sandbox, feed captured output back as a `tool` message; stop on a plain
answer or the cap; stream per-call progress to chat. Degrades to _run_simple when the
provider has no complete_with_tools or the model rejects tools.
- _exec_prefix/_exec_capture/_exec_tool: <engine> exec into docker/podman/multipass/local;
paths passed as positional args + content via stdin (no shell interpolation); combined
stdout+stderr byte-capped + time-bounded. run_shell is the only intentional shell.
- Guards: DESTRUCTIVE run_shell commands are blocked (not run — no human in the loop;
use simple + /ai confirm for destructive intent); MAX_COMMANDS budget per task.
- _run_in_sandbox dispatches native|simple; default harness flipped to native.
__main__.py: default harness native (self-degrades to simple, so safe).
Offline-tested: full write/run/read loop on the local backend; destructive block
(rm -rf never executed); ToolsUnsupported → simple fallback. Live Ollama wire
validation deferred to Phase 3 bench (daemon was down). py_compile clean.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
@@ -23,6 +23,11 @@ class Msg:
|
||||
content: str
|
||||
|
||||
|
||||
class ToolsUnsupported(RuntimeError):
|
||||
"""Raised by ``complete_with_tools`` when the backend model can't do function
|
||||
calling — the native harness catches it and degrades to the simple injector."""
|
||||
|
||||
|
||||
@runtime_checkable
|
||||
class Provider(Protocol):
|
||||
name: str
|
||||
@@ -56,6 +61,11 @@ class OllamaProvider:
|
||||
self.num_predict = num_predict
|
||||
self.num_thread = num_thread
|
||||
self.keep_alive = keep_alive
|
||||
# Tri-state tool-calling capability cache: None=unprobed, True/False once a
|
||||
# real /api/chat with `tools` either succeeds or is rejected by the model.
|
||||
# The native harness reads this to skip retrying tools on a model that
|
||||
# can't do them (and fall straight to the simple injector).
|
||||
self._tools_ok: bool | None = None
|
||||
|
||||
def _options(self) -> dict:
|
||||
opts = {"num_ctx": self.num_ctx, "num_predict": self.num_predict}
|
||||
@@ -105,6 +115,53 @@ class OllamaProvider:
|
||||
self._raise_for_status(r)
|
||||
return (r.json().get("message", {}).get("content") or "").strip()
|
||||
|
||||
def supports_tools(self) -> bool | None:
|
||||
"""Cached tool-calling capability: None until the first ``complete_with_tools``
|
||||
call has either succeeded or been rejected by the model."""
|
||||
return self._tools_ok
|
||||
|
||||
def complete_with_tools(
|
||||
self, system: str, messages: list[dict], tools: list[dict]
|
||||
) -> tuple[str, list[dict]]:
|
||||
"""One non-streaming ``/api/chat`` turn carrying a ``tools`` schema. Used by
|
||||
the native harness loop. ``messages`` are raw Ollama wire dicts (so the
|
||||
caller can round-trip assistant ``tool_calls`` and ``tool`` results across
|
||||
turns); ``system`` is prepended. Returns ``(text, tool_calls)`` where each
|
||||
call is ``{"name": str, "arguments": dict}``. Raises ``ToolsUnsupported`` if
|
||||
the model can't do function calling so the bridge can fall back to simple."""
|
||||
payload = {
|
||||
"model": self.model,
|
||||
"stream": False,
|
||||
"keep_alive": self.keep_alive,
|
||||
"options": self._options(),
|
||||
"tools": tools,
|
||||
"messages": [{"role": "system", "content": system}] + messages,
|
||||
}
|
||||
r = requests.post(f"{self.host}/api/chat", json=payload, timeout=self.timeout)
|
||||
if not r.ok:
|
||||
try:
|
||||
detail = (r.json().get("error") or "").strip()
|
||||
except ValueError:
|
||||
detail = (r.text or "").strip()
|
||||
if "does not support tools" in detail.lower():
|
||||
self._tools_ok = False
|
||||
raise ToolsUnsupported(detail or f"{self.model} does not support tools")
|
||||
self._raise_for_status(r)
|
||||
self._tools_ok = True
|
||||
msg = r.json().get("message", {}) or {}
|
||||
text = (msg.get("content") or "").strip()
|
||||
calls: list[dict] = []
|
||||
for tc in msg.get("tool_calls") or []:
|
||||
fn = tc.get("function") or {}
|
||||
args = fn.get("arguments")
|
||||
if isinstance(args, str):
|
||||
try:
|
||||
args = json.loads(args)
|
||||
except ValueError:
|
||||
args = {}
|
||||
calls.append({"name": fn.get("name", ""), "arguments": args or {}})
|
||||
return text, calls
|
||||
|
||||
def stream(self, system: str, messages: list[Msg]):
|
||||
"""Yield reply text incrementally as Ollama generates it. On CPU the
|
||||
perceived latency is TTFT, so streaming makes a slow reply feel live."""
|
||||
|
||||
Reference in New Issue
Block a user