Add sandbox template image with Claude configuration and plugins
build / build (push) Canceled after 0s

Carries CLAUDE.md, AGENTS.md, hooks and skills verbatim from the host, plus a
manifest of the 10 marketplaces and 17 plugins to reinstall at build time. The
plugin directories themselves are not committed: ~/.claude/plugins is 831 MB and
sits alongside credentials and transcripts, so the image is reproduced from the
manifest instead and the build needs no access to the host.

The Gitea registry is behind Cloudflare, which rejects request bodies over 100 MB
against a base image with a 325 MB layer, so the workflow pushes chunked through
regctl rather than docker push.

sbx v0.37.0 and v0.37.1 cannot consume the result: layers stacked on the base are
silently dropped (docker/sbx-releases#366). The image builds and pushes correctly
and is a no-op at runtime until that is fixed, so README points at
'ai:sbx setup' as the mechanism that works today.
This commit is contained in:
2026-07-31 08:18:30 -05:00
commit c89e4f0568
22 changed files with 2726 additions and 0 deletions
@@ -0,0 +1,98 @@
# Sliders, Not Checkboxes — classification reference
The framework models work as a small set of **feature categories**. A backlog
item is restated as the **need** it addresses (not the proposed solution), and
needs are value-ranked within each category. Source: the "Sliders, Not
Checkboxes" paper (CareEvolution).
## The need-restatement discipline
Before an item enters a category, restate it as the user/admin/developer
**need** — what they are trying to accomplish — and treat the proposed solution
as *one way* to address it.
- A doc says "build a self-serve provisioning UI." The **need** is "add/remove a
user without filing a support ticket." The UI is one solution; SCIM or a
declarative API are others. Title = the need; list the candidate solutions in
the description.
- Ranking reflects what addressing the need means for the people who hold it:
**how often** it's hit, **how acute** the pain, **how broadly** held. Not what
is easiest to build or most recently requested.
- Discussion, debate, and context that isn't a candidate item → do not create a
task. Capture only genuine candidate needs.
## The 8 feature categories
Classify each need into exactly one. Pick the category of the **primary value**;
note a close alternative in the description.
| Category (use this exact name) | What it covers | Whose need |
|---|---|---|
| **End-user experience** | User-facing product: features, flows, learnability, polish | People using the front-end to do their own tasks |
| **Admin & operator experience** | Configuration, governance, monitoring, operational tooling | Customer-side admins / IT staff deploying & overseeing the product |
| **Developer experience & adoption** | API design, SDKs, docs, sandbox, time-to-first-call; community, mindshare | Developers integrating against or building on the product |
| **Data, integrations & ecosystem** | Data quality, lineage, connectors, standards conformance (FHIR, HL7, OAuth, OpenAPI), marketplace, partners | Integrating systems and partner platforms |
| **Performance & scale** | Latency, throughput, behavior under load | Anyone depending on the product being fast and handling their volume |
| **Reliability & observability** | Uptime, fault tolerance, graceful degradation, recovery; logs, metrics, traces, alerting, runbooks | Customers needing it up; operators keeping it up |
| **Security, privacy & compliance** | Authn/authz, encryption, audit logging, certifications, regulatory posture, attack-surface reduction, vulnerability/CVE management, supply-chain integrity, runtime threat detection | Security & compliance teams; regulators; users trusting the platform |
| **Cost efficiency** | Cost per unit work — per record, per API call, per active user; compute/arch cost (e.g. ARM64/Graviton) | Internal P&L; customers indirectly via pricing |
### Classification hints / common overlaps
- **Hardening** (read-only rootfs, drop caps, minimal/chiseled base images, image
CVE scanning, SHA-pinning CI actions, dependency patch currency, supply-chain
inventory, runtime threat detection) → **Security, privacy & compliance**.
- **ARM64/Graviton migration, smaller images for cost** → **Cost efficiency**
(its perf side can be **Performance & scale** — pick the primary driver named
in the doc).
- **Perf vs cost** are often two sides of one coin — categorize by the value the
doc emphasizes (latency/throughput → Performance; $/unit → Cost).
- **A correctness/data-race bug surfaced by infra work** → **Reliability &
observability** (it's about the service being correct/up), not Security.
- **CI/CD integrity** (pinning actions, blocking tampering) → Security; **CI
ergonomics/speed for devs** → Developer experience. Pick by the value.
- **Runtime anomaly detection** (GuardDuty etc.) → Security (threat detection)
even though it touches observability.
## Value rank → priority
Vikunja priority is 1–5. Use **1–4** for value tiers (reserve 5 for true
DO-NOW). Rank within the category, by value of the need:
| Priority | Meaning |
|---|---|
| **4** | Rank-1 tier: most acute / frequent / broadly held; foundational. Everything else in the category is downstream of it. |
| **3** | High value, clearly worth doing, but not the keystone. |
| **2** | Medium: real need, smaller delta over the status quo. |
| **1** | Small / late-on-the-curve: refinement, niche, or largely satisfied by another item. |
Priority reflects **value, not build effort**. A cheap rank-1 outranks an
expensive rank-6.
## Urgency / returns labels
- **`acute`** — the need is frequent, painful, and/or broadly held *right now*.
Apply to the rank-1/keystone needs and anything externally forced (regulatory
deadline, audit finding).
- **`diminishing`** — the next investment in this area has tipped into
diminishing returns, or the need is largely satisfied by another item already
in the plan. Apply sparingly; it's a signal to deprioritize.
A task can have neither. Most have neither.
## Platform & deployment labels
- **Platform** = the product/platform the work belongs to (e.g. `Orchestrate`,
`HBNG`). One label.
- **Deployment / repo** = the specific service or repo (e.g. `Rosetta`,
`Hendrix`, `Insight`, `Kong`). One per task when granularity is per-deployment.
Categories are consistent **across** platforms and deployments (that's the point
— cross-deployment comparability). Platform/deployment are tags, never top-level
projects.
## Done / in-progress
- If the doc says an item is already shipped for a deployment, set `done: true`.
- Partial progress: leave `done: false` and set `percent` (e.g. 50) — note what
remains in the description.