← Back to the workbench

BUILD YOUR OWN WORKFLOWS

Your local AI lab, explained.

Ollama and the optional Colibri engine run models on this Windows PC, JORIS-DESKTOP. Debian hosts the Docker hub, public HTTPS entry point, existing Portal login and fallback page. Connected computers keep their project files and execute their own commands. OpenCode supplies file editing, shell commands and the coding-agent loop. This workbench manages tasks, comparisons, evidence and skills. JorisHoef Portal supplies your existing login. There is no additional account database and no paid inference provider.

Connect computers and configure the installation

Open Setup & portability for the four-step wizard: Portal, model computer, connection checks, and reuse elsewhere. Choose the destination system, storage folder and port to download one configured ZIP. The Docker hub is portable; Windows model engines, game detection and the tray still run natively.

Open Devices on any computer to download its companion. Run Connect.cmd on Windows or connect.sh on Linux/macOS, then approve its matching code in the portal. Open Projects to choose that computer, its existing folder, instructions and preferred model. A browser-only phone can control tasks; it cannot execute desktop tools. Other computers need the companion running and an internet connection to this domain. No incoming port needs to be opened on them.

1. Start with a small real task

Choose Chat & analysis for questions. Choose Coding agent for file changes and commands. Start with Qwen 3.5 9B, a clear task and a 15-minute budget. Project copy mode receives a separate copy of the selected project. Full PC access edits the paths you request directly, without registration. Closing the browser does not stop it. The model computer and the computer owning the project must stay awake. Interrupted tasks can be resumed explicitly.

The agent runs as the Windows account that started the service. Its workspace permissions are guardrails, not an operating-system sandbox. Use the existing portal administrator login. Project copy changes stay separate. Full PC access changes the original files you requested and uses the same limited Windows account; it does not automatically obtain administrator rights.

2. Compare models fairly

Comparison lab sends the same prompt to selected candidates sequentially. You can choose another model as a judge. It sees anonymized answers in two different orders to expose order sensitivity. Scores are model opinions; agreement does not prove correctness. Automatic judging currently applies to model-host tasks. For connected-device projects, start a separate review task with another model.

Measured baseline checks have explicit expected answers and report pass/fail. They cover four small tasks, not overall intelligence or coding quality. For code changes, inspect actual files and tests. Tokens per second measure generation; total time also includes loading and prompt processing. Use the same prompt and comparable context budgets. Ollama uses a fixed seed; Colibri does not implement per-request seed control, so repeat its trials. Colibri reports total time, startup time and time to first output; generation-only speed is left blank when the API does not provide it.

3. Give the agent a useful assignment

In this project, fix the stale-route bug after a terrain cell becomes blocked.
Read AGENTS.md and the terrain/navigation design first.
Add or use a regression check that fails before the fix and passes afterward.
Stay within the active milestone. Do not change the editor or renderer.
Save CHECKPOINT.md and report the exact test output.

A good task defines the outcome, boundaries and evidence. Larger models may handle a harder task differently, but they will use system RAM on this GPU and can be slower. First measure their behavior on your own tasks.

4. Make your own skill

Open Skills, choose New skill, and use a lowercase name with hyphens. The description tells the agent when to load it. The body contains the actual workflow. Save it; future agent jobs receive a copy.

---
name: unity-verification
description: Verify Unity C# changes using the project's installed editor and existing tests.
---

Read AGENTS.md and the project README.
Find ProjectSettings/ProjectVersion.txt and the existing test commands.
Use the matching installed Unity editor. Do not upgrade it.
Run the relevant checks and preserve logs.
If a check fails, identify the cause before changing code.
Report pass, fail or not run, with commands and log paths.
Do not call a source-only change compiled or playtested.

Keep one skill focused on one job. Include a small working script when a repeatable operation benefits from executable code. For a portable folder, use:

my-skill/
  SKILL.md
  scripts/       # optional tested helpers
  references/    # optional supporting documentation
  assets/        # optional templates

The portal stores approved SKILL.md text centrally and sends it to connected-device tasks. The native desktop library also remains at C:\AI-Workbench\data\skills. Supporting scripts and assets in native skill folders are currently copied only for native desktop tasks; remote companions receive the approved instruction text. Keep required helpers in the project itself until folder-bundle distribution is added.

5. Let an agent help create a skill

Use the included skill-author workflow in a coding task: “Create a skill for this recurring task, with precise triggers and one realistic example.” The agent writes its proposal inside its workspace. Review it, then paste the approved SKILL.md into the Skills editor or copy the folder into the library. The agent's proposal does not silently replace an approved skill.

6. What learning means here

Knowledge stores explicit notes: your preferences, confirmed project facts and lessons supported by tests. Skills store improved procedures. Checkpoints preserve work in progress. These are reusable context, not changes to model weights. Fine-tuning is a separate experiment requiring a dataset, training setup and evaluation; this installation does not pretend to do it automatically.

Model-reported thinking is available for models that emit it. It can help inspect behavior, but it may be incomplete or wrong. Important choices should also be explained briefly in the final answer, alongside evidence.

7. Tray icon, hosting and recovery

The AI tray icon appears when you sign in to Windows. Double-click it to open the Workbench. Right-click for status, hosting information, the local offline interface, and service start/stop controls. The status window explains what runs on this PC and on Debian. Exiting the tray keeps the AI service running; Stop AI service interrupts it. If Windows places the icon in the overflow area, open the arrow beside the clock. You can reopen it with C:\AI-Workbench\Open AI Tray.cmd.

Use C:\AI-Workbench\scripts\Start-Workbench.ps1 to start the runtime, dashboard and secure connection. Open-Workbench.ps1 opens the local offline interface. The public address uses the existing portal login. The service starts at Windows boot without requiring sign-in. Keep your PC awake for unattended tasks. If the model computer is off or asleep, the Debian hub stays available and reports that the model host is offline. Models and coding work wait for it to return. If the hub itself fails, nginx shows the fallback page. Logs are in C:\AI-Workbench\logs, tasks in data\jobs, and edits in workspaces.

Back up data, skills, configuration and any valuable workspaces. Models can be downloaded again. No automatic model or application upgrades run; pinned versions make comparisons repeatable.

8. Pause, gaming and resource limits

The large PAUSE ALL AI button appears on every portal page and in the tray. It holds models, agent commands and downloads; the portal stays available. Resume restores your saved limits. If a game is still running, inference stays paused. Files and sessions are retained, but an interrupted command may already have changed something, so the agent checks actual state before continuing.

PC resources offers Quiet, Adaptive and Full speed, a shared CPU ceiling, RAM headroom, GPU allowed or CPU only, and game-process names to include or ignore. Adaptive yields CPU time to other apps. RAM pressure pauses work; it is not a hard memory partition. GPU use has no percentage slider—CPU only keeps Ollama off the GPU. Colibri currently uses CPU/RAM. Games count while minimized, and the agent waits 20 seconds after they close.

Your selected behavior allows the GLM download at up to 24 MiB/s while gaming. Manual pause stops it too. Settings and partial downloads survive service restarts.

9. Full access and saved credentials

Choose Coding agent → Full PC access, then name the folder, server or command in your task. No per-folder or per-host registration is required. It can use installed command-line tools and SSH with the credentials you authorize. It runs under your limited Windows service account; interactive desktop control is a separate integration because the boot service has no signed-in desktop.

Save passwords or tokens in Credentials and refer to their short names in tasks. Windows encrypts the values on this PC. Names, labels and usernames are visible; secret values are not returned through the portal. The agent can supply a named value directly to a command. Full-access code under your Windows account can use this vault. Store ordinary facts in Knowledge and workflows in Skills.

10. Colibri experiments

Select Qwen 3.6 · 35B · Colibri in Chat & analysis or Comparison lab. This first integration uses the native Windows CPU/RAM engine, an 8K context, and verified int4-gs64 weights (23.0 GB). It supports text conversation and report-only review; native coding tools and image input are not enabled. Your original coding models still use OpenCode and Ollama.

Colibri starts only when selected and stops after 90 seconds idle. Switching to Ollama stops Colibri; switching to Colibri unloads Ollama models. The queue runs one workload at a time. A cold request includes model startup and prompt processing; this experiment has not been tuned for maximum speed.

GLM 5.2 is the large experiment: approximately 400 GiB of checked weights. Download progress appears in PC resources. Its native tool check waits until the complete download is verified and your resource settings allow work. Coding-agent use is enabled only after that check passes. A successful small fixture still needs to be followed by tests on your own real tasks.

OpenCode skill documentation · Ollama documentation

11. Work until verified

Choose Coding agent → When is this task finished? → Until finished · verify and repair. Add a required outcome and, where possible, a real verification command. Empty commands require your judgment. Project settings can save these checks for future tasks.

The controller runs checks separately, saves their output, and requests another repair attempt after failures. Add existing test files as protected paths so they cannot be quietly weakened. Tasks stop for blockers, repeated non-improving checks, overall limits, or your pause/stop controls. A session boundary saves progress and continues. No limit selection removes the overall time limit; attempt limits still apply.

Verified means the configured checks passed for the recorded file revision. Needs your review means human judgment or reviewer findings remain. After inspecting the saved result, use Accept this result after my review. This does not claim another automated test. Connected devices need companion 0.5.0. One model task remains active at a time in this release.