Skip to content

Known limitations

What to know before you rely on steer. None of these are bugs - they're the edges of what the plugin can guarantee, given the surfaces and tools it runs on. When in doubt, fall back to human review.

Where hooks fire (Claude Code vs. the chat surfaces)

Plugin hooks are a Claude Code lifecycle feature, and where they fire depends on the surface (validated June 2026). The Claude Desktop app has three tabs - Chat, Cowork, and Code - and they don't behave the same:

  • Claude Code - the CLI, the IDE extensions (VS Code / JetBrains), and the Claude Desktop Code tab - runs hooks fully. The always-on rules inject, the PreToolUse gates run, skills and MCP work. This is the supported path.
  • Cowork (the Cowork tab) is the one chat-family surface where hooks and sub-agents run - Anthropic's docs state "hooks and sub-agents run only in Cowork." Plugin-scoped SessionStart hooks had bugs earlier in 2026 (since closed); reconfirm on your build before relying on auto-injected rules there. Cowork is best-effort and PO/knowledge-work only - it's a no-install sandbox (see below), so engineering work belongs in Claude Code, not here.
  • The Claude Desktop Chat tab and claude.ai web chat do NOT run hooks - they show as grayed out. Plugins install and skills work, but the always-on rules are not auto-injected and the PreToolUse gates don't run.

On the no-hooks surfaces (Chat tab, web chat) - and as a fallback anywhere the rules didn't load - run /steer:standards at the start of the session to load the rules by hand, and rely on human review where the gates would have fired. The rules are the one that matters: they're what make Claude follow the standards; the PreToolUse spec-first and issue-first nudges are only advisory even when they do fire (see below), so losing them matters less. See Installation and the Hooks reference.

Knowledge-work mode (non-code folders, e.g. Claude Cowork)

When a session opens a folder that is confidently not a code project - no git work tree and no code/config markers nearby - steer injects a lean, PO-relevant ruleset instead of the full engineering manual. This is the typical Claude Cowork case: a product owner opens a connected folder of specs/docs. In that mode only the unmarked rules inject - since the 6.6 rule diet that is six of eighteen: 00-router, 03-output, 05-roles, 30-spec, 60-high-risk and 61-gates. Every rule carrying any inject-when marker is skipped, whichever token it is: the code-project set (35-tracker, 40-testing, 45-delivery, 50-done, 80-change-class, 85-practices), the org pack (10-stack, 15-commands, 12-stack-infra), the OpenSpec backend, issue-first, and the opt-in loop rule. The skip is intentional - it reclaims context budget and cuts noise - and orient-session confirms in plain language that the standards are active.

One of the injected rules reads as code-specific and is always-on anyway, by design: high-risk areas (60-high-risk, which since 6.6 also carries the secrets standard as a section). Three of the five other surviving rules cross-reference it - 05-roles for the PO guardrails, 30-spec for what needs an ADR, and 61-gates for what is never promptable - so dropping it would break them. Run mise run rules:preview -- --knowledge for the authoritative inject/skip table; the numbers here are from it, not from memory.

The classification is fail-safe: a git repo, any code/config marker, or any uncertainty resolves to full code mode - steer never silently drops a rule from a real code project. /spec is deliberately not treated as a code marker, since a knowledge folder is exactly where a spec spine may live. The limitation to know: in a non-git folder steer cannot detect a code project that carries no on-disk markers, so a marker-less code checkout opened without git would get the lean set - add a mise.toml/package.json (or open it as a git repo) to get full rules.

Claude Cowork's sandbox: no installs, connector-only GitHub

Claude Cowork runs in an Anthropic-managed, sandboxed Linux VM (OS-level isolation via bubblewrap/seatbelt), not a normal dev machine. Claude can read, write, and run scripts inside the connected folder, but the sandbox's filesystem and network are locked down: in practice you cannot install system tooling - docker, mise, language toolchains, or the gh CLI - the way you can in a Claude Code CLI session (validated June 2026). Treat Cowork as a no-install surface. This is an environment boundary, not a steer bug.

Two consequences follow, and together they explain why "the GitHub connector isn't working" in Cowork even though it works in the CLI.

1. The plugin's .mcp.json is a Claude Code mechanism - Cowork doesn't use it. MCP config is not shared across surfaces: Cowork wires MCP through its own Connectors, not the plugin-shipped plugins/steer/.mcp.json that the CLI reads. So of the two servers steer ships, one does not survive Cowork - and the no-install rule separately takes out document conversion:

  • github authenticates with Authorization: Bearer ${user_config.github_pat}, resolved from the plugin's user config (prompted at install, held in the macOS Keychain or ~/.claude/.credentials.json). Cowork reads neither the CLI .mcp.json nor that config store, so the plugin's GitHub server appears to "try to connect like Claude Code" and fails to authenticate. Do not rely on it in Cowork.
  • context7 is a plain hosted HTTP endpoint with no token, so it is the one that does work if the surface routes it - nothing to install, no shell secret.
  • Office-document conversion is not a server at all - it is the mise run convert:doc task (uvx --from 'markitdown[all]' markitdown). It needs uv/Python, which the sandbox cannot install, so /steer:spec intake drops to its manual floor in Cowork: it commits the binary and stops before diffing rather than fabricating an extraction. Convert the document elsewhere, or work in the CLI.

2. GitHub on Cowork = the built-in connector, not the plugin server. To do issue work in Cowork, enable the built-in GitHub connector (Cowork -> Customize -> Connectors), which Anthropic manages via OAuth and runs outside the bash sandbox. Once it's on, /steer:tracker-sync's MCP-first probe finds the repo-scoped issue tools (list / get / create / comment / label / transition) and /steer:work issues triage works - Cowork can triage GitHub issues. Caveats:

  • It is repo-scoped only. Org/team-level reads come back empty by design, so anything needing org config - Issue Types, and the org-level native issue fields (Priority/Effort/dates) field-set writes - may be unavailable and will degrade to the steer:kind marker / a human follow-up rather than fail loudly. Plain triage (read, classify, label, comment, set the steer:state marker, link issues) does not need org scope and works.
  • The gh-CLI fallback is unavailable (can't install gh), so when the built-in connector is off there is no automated path - only the manual floor.

Net: in Cowork, do issue triage through the built-in connector; for the install-dependent parts of steer (docker/mise builds, the convert:doc task, gh-CLI flows) use the Claude Code CLI or the Desktop Code tab, which share the full engine.

Headless vs. interactive runs

The plugin's gates assume an interactive human is present to approve specs and merges (pushing the branch and opening the PR are autonomous; only a solo-trunk repo's graduation gate (check-bash-actions.sh) pauses a push, and only when a local signal stands). In headless or scheduled (cron) runs there is no human at the gate, and interactively-authenticated MCP servers may be absent, so the MCP-first tracker path can silently fall back. Don't run the gated workflows unattended and expect the approvals to happen - /steer:loop is the sanctioned scheduled/headless path precisely because it stops at those gates (it never merges).

GitHub auth / gh

Anything that touches the tracker or GitHub needs an authenticated path. Tracker I/O routes MCP-first -> gh fallback -> manual floor (see the tracker-sync skill), so without an MCP tracker tool and without an authenticated gh, these operations drop to the manual floor:

  • creating and transitioning issues,
  • opening and updating PRs (gh pr create),
  • syncing lifecycle state.

Check with gh auth status. Skills never hit gh/MCP for issues directly - they go through /steer:tracker-sync, so the fallback is consistent, but the underlying capability still has to be there.

GitHub Projects automation

steer does not automate or manage a GitHub Projects board. The backlog is issue-first / local-first; triage lives in issues and the /spec spine. Priority, effort, and start/target dates are native GitHub issue fields on the issue (not labels, not Project-item fields) - steer reads them and escalate-only auto-sets Priority.

What steer does guarantee is that issues are Projects v2-compatible by construction: it sets the native attributes a board or roadmap reads - Issue Type, labels, assignees, milestone (/steer:tracker-sync set-milestone), native parent/sub-issue links, and the native issue fields above - so you can build an (org-level) board or roadmap on top without the plugin owning it. An epic (a parent tracking issue grouping features as sub-issues) makes the full Epic -> Feature -> Task tree board-visible by construction, so a Projects v2 Hierarchy view renders it with no extra machinery; Type=Epic is used only when the org enables that type, otherwise the epic carries the steer:kind=epic marker with its Type left unset. Only Project item custom fields (Status, iteration, size) live Project-side and are never written into the issue; steer:state stays canonical in the body and is mirrored at most one-directionally by a Project Status field.

The native-field vs Projects-column trap

When a Project v2 board surfaces the native issue fields (Priority, Effort, dates), they appear as single-select columns that look identical to genuine Project custom fields (Size, Iteration) but are API-locked. Every Projects write path rejects them - updateProjectV2Field and gh project item-edit return Only custom fields can be updated. Fields derived from issues or pull requests must be updated through their respective APIs - and every Projects read path reports options: [], so there is no option id to set through the Project at all, even though the UI shows Urgent/High/Medium/Low. Set these on the native issue field via /steer:tracker-sync field-set, never the Projects API. The reverse holds for a genuine Project custom field (Size, Iteration): it is not a native issue field, so it is edited with gh project item-edit and field-set will not find it. To populate a chosen Priority/Effort value (PO seeding, not the escalate-only floor), /steer:work issues triage and board route the request straight to field-set. field-set writes the native field through GraphQL setIssueFieldValue (or the equivalent REST issue-field-values endpoint) - not a GraphQL-only path, despite the Projects columns being read-locked. The REST fallback is a POST to /repos/{owner}/{repo}/issues/{n}/issue-field-values; never PUT that endpoint for a single field - PUT replaces all of the issue's field values, silently clearing the Priority/Effort/dates you did not pass.

Context window, compaction, and sessions

steer cannot manage your context window for you, and that is a hard Claude Code boundary, not a plugin gap: no hook or environment variable exposes the token count or how full the window is, and neither a hook nor the model can trigger /compact or start a new session - only you can. So steer will never silently compact or "switch you to a fresh session" when a long run fills the window.

What it does instead (the context-hygiene standard; the router carries the two always-on lines, full prose via /steer:reference context-hygiene):

  • Delegates heavy, multi-phase, or search-heavy runs to subagents, which get a fresh context window by construction and return only the result - so the heavy intermediate context never lands in your main session. Read/search/summarize fan-out delegations can run on a cheaper Sonnet-tier model at low effort; reviewer/verify/judge delegations stay on the session model.
  • Keeps durable run-state and task constraints in files (/spec/**, sidecars), which survive compaction and a fresh session where chat history does not. The SessionStart hook also re-injects the rules after a compact.
  • Only when the thread is genuinely overloaded does it recommend you /compact or start a fresh session - with a pre-composed hand-off - saying plainly that acting is your call, not something it can do.

One compaction behaviour is worth knowing about, because it shapes how steer's skills are written. An invoked skill's content stays in the conversation for the rest of the session, and when auto-compaction fires Claude Code re-attaches the most recent invocation of each skill - but keeps only the first ~5,000 tokens of each, with re-attached skills sharing a combined budget. A skill whose body runs past that cap silently loses its tail mid-run, and the tail is typically where guardrails sit.

steer works with that rather than against it: guardrails, coupling rules, and output contracts live near the top of every SKILL.md, and per-mode or per-phase procedure lives in sibling files the skill reads just-in-time for the path it is actually executing. So on a long /steer:work or /steer:audit run that compacts, the safety rules survive by construction, and the step-by-step detail is simply re-read when needed. A CI gate caps each SKILL.md body so this cannot regress (see Configuration). Invoking many skills in one session can still push older ones out of the shared re-attach budget entirely - a fresh session is the fix there.

What the hooks do (and don't) enforce

Even when hooks fire, only one of them hard-blocks an action, and one more raises a prompt. Be honest about the tiers:

  • SessionStart -> inject-standards.sh injects the rules. Real and load-bearing.
  • SessionStart -> session-checks.sh is otherwise read-only, with one exception worth knowing: check-worktree-trust.sh runs mise trust in a linked worktree to inherit the primary checkout's decision, so a fresh worktree's mise run ... doesn't fail on trust rather than on the task. It never creates trust - an untrusted primary checkout leaves it untouched and says so, keeping that first decision yours. Everything else the roster does is report-only. The same script also runs on CwdChanged, so a worktree entered mid-session gets the same inherited trust.
  • SessionEnd / WorktreeRemove are the only hooks that act on your Docker state, and only ever inside a linked worktree: SessionEnd stops that worktree's services (volumes kept), WorktreeRemove runs the full docker:clean because the checkout is being deleted. A plain checkout is never touched, and setting STEER_NO_WORKTREE_TEARDOWN to any non-empty value turns both off. Neither can block - no decision control - and both discard their JSON output fields.
  • The SessionEnd teardown is best-effort and will often not complete. SessionEnd hooks share a 1.5-second budget, and a timeout declared by a plugin does not raise it - only one in your own settings file does. steer's declared "timeout": 60 is therefore inert, and mise tasks ls plus mise run ... docker:down frequently exceeds 1.5s, in which case the hook is cancelled mid-run. Do not rely on exiting a session to free a worktree's ports. Raise the budget yourself if you want it to fit: CLAUDE_CODE_SESSIONEND_HOOKS_TIMEOUT_MS=5000 claude. WorktreeRemove takes the ordinary command-hook timeout and is the dependable half.
  • WorktreeRemove fires only for worktrees Claude Code deletes. Orca, Conductor and a plain git worktree remove delete theirs without it, so the stack keeps running. Trust inheritance, per-worktree ports and the SessionEnd stop still apply. /steer:setup worktrees installs the tool's own teardown - for Orca, an orca.yaml archive hook, which needs Orca's repository hook policy to allow shared scripts and which orca worktree rm skips without --run-hooks - and sweeps the stacks already orphaned.
  • The CwdChanged trust notices do not reach you. check-worktree-trust.sh applies mise trust fine on that path, but it writes its human-facing notices to stdout, and CwdChanged stdout goes to the debug log rather than the transcript. So a worktree entered mid-session whose primary checkout is itself untrusted gets no visible prompt - you will see the mise run ... trust error instead.
  • PreToolUse -> check-write-nudges.sh (the spec/scaffold + issue-first dimensions) is an advisory nudge that lets the write proceed. It is explicitly "a nudge, not a gate," fails open on any ambiguity, and the issue-first dimension only fires in GitHub-tracked repos. The spine reminder fires once per session about the missing /spec spine, but the scaffold reminder is sticky - it re-fires on each new feature file while the repo has no root mise.toml, since the bundled scaffold is product-independent and shouldn't be silently skipped (it still never blocks).
  • PreToolUse -> check-version-pins.sh is the only hard deny - it blocks image/runtime pins below the supported floor.
  • PreToolUse -> check-bash-actions.sh is an ask, never a deny: in a solo-trunk repo that has outgrown pre-MVP, the first git push to trunk of a session raises a permission prompt (approving it pushes anyway; the gate clears by graduating via /steer:setup protect). The same script also carries an advisory issue-create guard. Once-per-session-and-repo - later pushes downgrade to a note.
  • PostToolUse -> format-on-write.sh formats a file after it is written. Cosmetic and non-blocking; it never rejects or reverts the write.
  • PostToolUse -> check-comment-density.sh notes a source or config file whose comment lines exceed a fifth of its non-blank lines (rule 03-output § Code comments). A once-per-file-per-session notice, not a gate - the write already happened and is never reverted.
  • Stop -> reconcile-issue-first.sh reports, at end of turn, work that never got an issue. A report, not a gate - it cannot undo anything.
  • The merge gate is not a hook at all - it's a rule Claude follows (45-delivery, 00-router § You are not the gate). Nothing technically prevents a PR merge; a human reviewer is the real backstop.

When hooks fail or don't run

If a hook errors or simply doesn't fire (see Cowork/Desktop above), the session keeps working - but the rules aren't injected and the PreToolUse hooks don't fire. The plugin's hooks fail open by design (any ambiguity -> allow), and there is no automatic retry. Mitigation:

  • Load the rules manually with /steer:standards.
  • Rely on human review at every decision gate.
  • On managed surfaces, confirm rules loaded (the session should reflect the standards) before trusting the gates.

One failure mode is not fail-open - run /steer:setup doctor first

If every steer script dies at once with syntax error near unexpected token $'{\r', that is not a hook failing open - it is a CRLF-corrupted install, and a CRLF shell script does not warn, it fails to parse. The hooks share hooks/lib/*.sh, so one bad checkout takes out the whole set simultaneously (the v5.0.0 fault). It is not a case for /steer:report: it has a local, immediate answer. Run /steer:setup doctor - its §0 plugin-integrity check greps the installed hooks/ and scripts/ for CR before anything else and reports it as an install fault, with the repair. See Windows setup -> Line endings.

steer ships the real skill bodies to .agents/skills/ so Copilot, Cursor, Gemini CLI and Codex read the same procedures Claude Code does. Three things are rewritten on the way out, because they do not travel (the full table is in Copilot support); one of them is known to be incomplete, and it is recorded here rather than left to be rediscovered.

References to shared files - anything a skill points at outside its own directory - are rewritten to URLs on this public repo instead of being vendored into every consumer repo. Two defects in that rewrite are open:

  • The URLs point at GitHub's HTML blob/ view, so fetching one returns a rendered web page rather than the file's bytes. raw.githubusercontent.com is the form that returns content - except for a directory URL, which has no raw equivalent at all.
  • The rewrite is applied unconditionally, including inside runnable command lines, so the generated tree contains sh "https://..." and python3 "https://..." invocations that cannot execute on any surface.

What this means in practice. The rewrite is unconditional, so this is systemic rather than a short list of affected skills: any step that depends on reading a shared file or running a shared script does not work on a non-Claude surface. Some skills lose only a link; others lose their procedure, because fetching the file is the step - steer-standards is the extreme case: its one actionable step is a directory URL to read the rules from, and a directory URL has no raw. equivalent to fall back on. Treat the portable tree as reliable for guidance and unreliable for any instruction that reaches outside the skill's own directory.

On Claude Code none of this applies: it reads the skills from the installed plugin, where every path resolves normally.

Neither is fixed on this release. The first is a one-line change nobody has made; fixing the second is a design question - vendor the few helper scripts, fetch them to a temp file first, or drop those command blocks from the portable copy - so it is deliberately unpatched rather than guessed at. scripts/gen_agent_skills.py carries the same note in its module docstring.