跳到主要内容
知仓学习社ZHICANG

code-task

PREFERRED way to change code in a REAL repository: fix a GitHub issue, fix a bug, add/implement a function or feature, or make any edit to a project…

不碰外部(只输出文字)无严重或高危命中TokenRhythm/opensquilla

它会碰到什么

扫了多少1 个文本文件,11 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

code-task

Solve a real-repository coding task end to end: clone the repo to a

disposable working directory, run an OpenSquilla agent to make the change on

a task branch, then independently verify it with a red→green→regression

loop. Host mode (no Docker) in v1.

Use this — do not hand-edit the repo yourself

When the user asks to fix/add/implement/change code in a repository they

name by path or URL, route it through opensquilla code-task solve — even

if the change looks small enough to do by hand. Editing the files yourself

in this session is not equivalent: it skips the disposable clone, the

task branch, and (most importantly) the runner-verified red→green→regression

proof, so neither you nor the user gets evidence the change actually works.

Answer inline (no code-task) ONLY for truly trivial one-liners, pseudocode, or

conceptual / non-deterministic questions. For self-contained TESTABLE code from

scratch with no repo named, use `code-task solve --task "..." --verification-mode

scratch` (no --repo): it writes the code plus a test and verifies it green-only

(no red/regression -- there is nothing pre-existing to regress). When a real repo

is named, prefer code-task red-green.

Translating the user's request

The user speaks naturally ("fix issue 412 in github.com/acme/widgets",

"add CSV BOM support to my project at ~/code/foo"). Map that to the command:

opensquilla code-task solve --repo <url-or-path> ( --issue N | --task "<text>" | --task-file <path> ) [--yes]

> Invocation: do NOT assume a bare opensquilla (or bare python) is on PATH —

> the gateway commonly runs from an absolute interpreter path. When coding mode

> is active it injects the EXACT, resolved, PATH-independent command: use ONLY

> that. Otherwise invoke via an ABSOLUTE interpreter, e.g.

> /abs/path/python -P -m opensquilla.cli.main code-task solve .... Never

> pip install OpenSquilla or run an installer to "get" the command; if it

> cannot be run, stop and report the environment is broken.

  • A GitHub issue--issue N (needs gh; see below).
  • A short request in the message--task "<their request>".
  • A long spec, or pasted from Jira/GitLab/内网 → save it to a file and

use --task-file <path>.

  • Pass --repo for changes to an EXISTING repo. Omit it for

--verification-mode scratch (from-scratch code) and for a from-scratch

--verification-mode build (a brand-new app) — both scaffold their own repo.

If the user is already in a local checkout, use that path; otherwise the URL.

  • Pass --yes to skip the interactive trusted-host confirmation (you are

acting on the user's behalf), but only after the safety check below.

Before you run — two checks

  1. Trusted repo: code-task runs an agent on the host that may install

dependencies and execute the repo's code. It is NOT a sandbox. Only run

it against repositories the user trusts. If the repo's provenance is

unclear, ask first.

  1. Enough information: you must be able to state the expected behavior

change ("what is wrong/missing now, what should be true after"). If the

request is too vague to write an acceptance test for, ask the user to

clarify BEFORE running — do not burn a run on a guess.

  • Build-from-scratch (--verification-mode build) has no acceptance

test. Decide by whether you know WHAT THE APP SHOULD DO, not just its kind.

If the request is only a broad app type/goal with no concrete features,

target user, or scope (e.g. "make me an English-learning app", "a drawing

app"), ask 1-2 focused questions (core features/screens and who it's for),

then STOP — do not run code-task until answered. If it already names

concrete features, scope, or target users, do NOT ask — build it with

sensible defaults and state your assumptions. Never ask about

platform/framework/styling. At most 2 questions; never interrogate.

GitHub issue mode needs gh

--issue shells out to the GitHub CLI (gh). If gh is missing or not

authenticated, tell the user to gh auth login, or fall back: have them

paste the issue text and use --task / --task-file instead. The issue

body AND comments are pulled in (comments often hold the repro steps).

While it runs — watch the run dir, not the source repo

code-task clones the --repo source into an isolated run directory and does

all its work there. The **source repo stays empty until a run finishes and

VERIFIES**, at which point (build mode, local source) the change is committed

back. Therefore:

  • Do NOT judge progress by the source repo's contents, and do NOT conclude the

run is "stuck" because the source still looks empty — that is expected.

  • A run takes several minutes. Let it finish: process(action="wait") on the

background session. Do NOT kill it, do NOT "clean and retry", and do NOT

launch the same task again while one is still running.

  • The run prints its run directory on startup and writes a live

<run_dir>/status.json (phase = preparing → agent_running → collecting_change

→ verifying → completed). Watch that if you want progress.

  • Decide success only from the returned result state and

build.installer_path (which points into the run dir, not the source).

Reading the result

--json prints a result object; key fields:

  • state: verified (acceptance test went red→green, no regressions),

already_satisfied (the behavior already held on the base commit),

not_testable (work done but not expressible as a test),

environment_blocked (could not build/test the repo),

invalid_acceptance_test (agent produced no valid verification manifest),

failed (acceptance not green or a regression appeared).

  • branch, commits, files_changed, diffstat, patch_path.
  • acceptance: each test with beforeafter (e.g. failpass).
  • regression: existing-suite result and new_failures.
  • assumptions: surface these to the user — a wrong assumption means a

wrong fix.

  • usage: cost / tokens (aggregated across internal retry attempts).
  • attempts / max_attempts / retry_exhausted: code-task RETRIES internally

when its own verification fails -- it re-runs the agent on the SAME prepared

repo (no re-clone, no re-explore) with the concrete failure fed back, up to

max_attempts. A returned result is FINAL across those internal attempts.

  • relaunch_recommended is always false and final_failure_reason explains a

failure: do NOT re-launch the same task yourself on a failed result -- the

internal retries are already exhausted. Surface the failure to the user.

What to tell the user

  1. Up front: cloning + dependency install + the agent loop can take several

minutes; you'll report when done.

  1. After: report state, what changed (diffstat), the acceptance red→green

evidence, any assumptions, the cost, and where the branch/diff lives.

  1. On failed / environment_blocked: quote error / final_failure_reason

and point at the agent_stdout.log. Do NOT relaunch the same task yourself --

code-task already retried internally (see retry_exhausted).

Constraints

  • Runs on the gateway host — git, the toolchain, and disk all come from

there. Works the same from TUI, Web UI, or any channel.

  • v1 is host-only and always clones fresh (no --in-place). For untrusted

repositories, a Docker-isolated backend is planned but not in v1.

Verification modes

code-task solve defaults to --verification-mode red-green: the agent writes acceptance tests, the runner proves red on the base and green on the change, then runs regression.

For building an app or UI from scratch (e.g. an Electron + Vite + React desktop app) there is no red->green test loop. Use --verification-mode build: the runner owns a fixed checklist (npm ci -> npm run build -> npx electron-builder for the HOST OS, with the target pinned — --mac dmg -> .dmg, --win nsis -> .exe, --linux AppImage -> .AppImage) and state=verified means the app actually builds and packages into an installer. Each OS only builds its own installer, so run on each OS (or a CI matrix) to collect all three. The result carries verification_kind=build and build.installer_path(s). Preview/launch is intentionally out of scope (no GUI is run).

For self-contained, testable code when the user has not named a repo, use

--verification-mode scratch with --task or --task-file and no --repo.

The runner creates an empty git repo, asks the agent to write code plus pytest

coverage, and independently reruns the declared acceptance command. This mode is

green-only and returns verification_kind=scratch.

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 7,018
本站分层T1
该仓技能数68
原文件路径src/opensquilla/skills/bundled/code-task/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 68 个技能