Test Execution Handbook
Neemias
1. Purpose
This handbook converts test strategy into execution rules for contributors.
Use this with:
mtp.md../../RTM.csv../architecture/srs.md
2. Scope
Applies to:
- unit, integration, and system test execution
- accessibility and compliance verification steps
- release-gate evidence requirements
3. Requirement Traceability Rule
Every behavior-changing PR should reference affected requirement IDs and test IDs.
Minimum PR traceability:
- Requirement IDs from
../architecture/srs.md - Test case IDs in PR summary
../../RTM.csvupdates when mappings change
4. Test ID Naming
Use stable IDs:
TC-RBAC-###TC-ATT-###TC-STU-###TC-SYNC-###TC-CONFLICT-###TC-TTL-###TC-A11Y-###TC-SEC-###TC-COMP-###TC-PERF-###
5. Planned Test Structure for Implementation
When code scaffold is introduced, follow this structure:
tests/unit/tests/integration/tests/system/tests/accessibility/tests/security/tests/fixtures/
Note:
Tooling commands are intentionally not fixed in this document until the implementation stack is finalized.
6. Execution Order by Risk
Run tests in this order when validating release candidates:
- Authentication and role gating (
SCR-001,SCR-002,FR-001toFR-008) - Offline queueing and sync conflict rules (
FR-014,FR-015,SR-001toSR-005) - Data integrity and audit retention (
DR-001toDR-005) - Deletion justification and admin-only constraints (
FR-007,SCR-007) - Accessibility and UX confirmation (
NFR-001,FR-021,FR-022) - Localization and defaults (
FR-019,FR-020)
7. Evidence Requirements
Each executed test cycle should capture:
- timestamp and environment profile
- test ID and requirement mapping
- pass or fail result
- logs and screenshots for failures
- defect ticket reference for failures
Mandatory artifact groups:
- conflict resolution evidence
- session expiry and re-auth evidence
- deletion justification evidence
- accessibility quick-audit evidence
8. Defect Triage Rules
Severity follows mtp.md:
- Critical: security breach, data loss, conflict rule break, compliance breach
- High: blocked business flow, incorrect sync outcome
- Medium: recoverable behavior mismatch
- Low: cosmetic and wording issues
Gate policy:
- no open Critical defects for release
- no open High defects in security, sync, or data integrity domains
9. Manual Accessibility Checklist (Minimum)
For each major flow (login, attendance, students, deletion, reports):
- full keyboard traversal
- visible focus indicator
- semantic labels for controls
- error association to fields
- non-color-only status cues
- screen-reader announcement for success and failure states
10. Session and Time-Travel Testing
Session expiry tests should use deterministic time controls where available.
Requirements:
- do not wait real 24 hours in tests
- advance clock abstraction in test harness
- verify unsynced queue remains intact across expiry
- verify protected submission blocked until re-authentication
11. Release Candidate Checklist
Before release candidate signoff:
- requirement coverage report generated against
../../RTM.csv - priority suites in
mtp.mdexecuted - unresolved risks documented
- accessibility checklist completed
- decision-impacting findings reflected in
../architecture/adr/when needed
12. Local E2E runbook
The Playwright suites run against a local stack and are deliberately not part of pnpm test:all (#750): the umbrella covers the unit tiers plus the Worker integration suite, while E2E needs a browser and (for the online suite) a Worker process.
Prerequisites
pnpm install
pnpm --dir app exec playwright install chromium # add --with-deps on LinuxNo env exports are needed: each config sets its own VITE_* values, and Playwright boots the servers it needs as webServer entries (reuseExistingServer locally, so an already-running server is reused). Ports used: 5173 (Vite), 8788 (wrangler dev, online suite only).
Tiers and commands
| Tier | Command | Stack | Specs |
|---|---|---|---|
| Smoke (what CI runs) | pnpm --dir app test:e2e:smoke | Vite dev on :5173, VITE_BACKEND_URL="" (offline island), demo seed, 1 worker, no retries, chromium | 9 specs / 18 tests (~20 s locally) |
Safari (nightly CI, safari.yml) | pnpm --dir app test:e2e:safari | identical to Smoke but on WebKit — needs playwright install --with-deps webkit (its Linux system libs differ from chromium's) | 11 specs / 22 tests: the 9 smoke specs plus 08-responsive.spec.ts and 11-accessibility.spec.ts (see below) (~29 s on an idle mac, ~3.5 min on the runner) |
| Full offline | pnpm --dir app test:e2e | same stack, 3 workers locally | all app/e2e/specs/*.spec.ts |
| Online (what CI runs) | pnpm --dir app test:e2e:online | wrangler dev on :8788 with a fresh local D1 (app/e2e/scripts/e2e-online-worker.sh) + Vite with VITE_BACKEND_URL=http://localhost:8788, VITE_ONLINE_ONLY=true, 1 worker, 1 retry | 09-modo-facil-checkout.spec.ts, 12-modo-facil-notificacoes.spec.ts, 13-modo-facil-cadastro.spec.ts, 14-modo-facil-qr-checkin.spec.ts |
Every test:e2e* script pins --project=chromium on purpose — playwright.config.ts also declares a webkit project (the PWA's primary browser: spec 14 calls iOS/Safari the fila principal), and both the smoke and online configs spread baseConfig. Without the pin, adding that project would silently run every suite on both engines and roughly double the e2e lane. To run Safari locally use test:e2e:safari; to run everything on both engines, pass --project explicitly. Note that desktop WebKit is not iOS Safari — it catches JS/CSS engine divergence, not the iOS install flow or storage eviction, which the manual device checklist owns (docs/quality/modo-facil-device-smoke.md; the canonical item list is the header of spec 14).
The Safari tier is the smoke subset PLUS two specs that no other tier runs (playwright.safari.config.ts imports SMOKE_SPECS from the smoke config and appends them, so the two lists cannot drift):
08-responsive.spec.ts— layout at 375 px (phone) and 1024 px (tablet). This is a phone-first PWA, so a 375 px regression is user-facing; nothing else in CI covers it.11-accessibility.spec.ts— axe-core over /login and /dashboard, asserting zero WCAG 2.1 AA violations. Its own header records that Students/Attendance/Classes are excluded forcolor-contrastviolations — those three pages are still ungated. Closing that gap is a separate, known task; do not read a green accessibility run as "the app is AA compliant".
Both are deliberately nightly-only: they are "deeper check" work, not the fast pre-push net, and they never ran in any CI before 2026-09-15 (both verified green on chromium and webkit first).
Notes:
02-attendance.spec.tstargets a removed UI and is intentionally outside the smoke subset — rewrite or delete it before adding it back (playwright.smoke.config.tsheader).- If Playwright reports
Executable doesn't exist at .../chromium_headless_shell-<build>/..., the cached browser is older than the pinned Playwright version — rerunpnpm --dir app exec playwright install chromium. - The online suite needs
workers/.dev.vars(dev admin token +SEED_ENABLED); itsglobalSetupseeds viaPOST /api/v1/_seed(a409 ALREADY_SEEDEDis fine). The file is gitignored, so CI synthesizes it in thee2ejob. - The smoke subset plus the online suite are the CI safety net (both run in the
e2ejob);pnpm test:allnever runs any Playwright spec. Both suites are skipped when the change is docs-only (if: steps.changed.outputs.code == 'true'in thee2ejob ofci.yml). - The
e2ejob starts by freeing the E2E ports withapp/e2e/scripts/free-e2e-ports.sh(:5173Vite,:8788wrangler dev). A job that is cancelled or hitstimeout-minuteson the shared self-hosted runner leaves its servers alive, and the next run fails at webServer startup —http://localhost:5173 is already used(run 34787497864) — before a single test runs, and Playwright's retries never cover webServer boot. The script kills listener + supervisor in a loop (a plain kill only makeswrangler devrestart itsworkerdchild) and exits 1 if a port still responds after 20 s, rather than let a stale server silently serve old code. Use it locally too beforetest:e2e:online: withreuseExistingServer: truea stale worker is otherwise silently reused and the fresh-D1 wipe is skipped. - The online specs need the real Worker (D1 + R2), so they cannot move into the offline smoke subset. Because they all log in from one IP, they exceed the login ceiling of 5/60 s; the local E2E Worker is booted with
--var RATE_LIMIT_E2E_MAX:500(ignored whenENVIRONMENT=production— seeworkers/src/middleware/rate-limit.ts). The per-route ceiling is not just about logins: the screens the specs drive pollGET /api/v1/studentshard, and a fast run fired 54 requests in a single 60 s window (read the counts straight from the localrate_limitsD1 table), which is why the earlier ceiling of 50 produced a deterministic429 RATE_LIMIT_EXCEEDEDfor spec 14 while the same suite passed on a slower runner. If a spec gets a 429, raise the ceiling — never add a login-driven spec's traffic to production's real limit. - Diagnosing an online-suite failure: read the artifact, not the spec names. The Worker on
:8788is the suite's infrastructure dependency, and its death mid-run does not look like infrastructure to Playwright: on CI run 34853673151wrangler devcrashed two minutes in and the three specs that followed failed asNenhuma opção disponível para a turma (seed do D1 local vazio?)andconnect ECONNREFUSED ::1:8788— zero bugs, three "failures". So:- The
e2ejob uploadsapp/test-results/as thee2e-online-diagnosticsartifact on every run (if: always()), which includes/tmp/wrangler-e2e.logand/tmp/wrangler-e2e-migrations.logcopied intotest-results/wrangler/— that is where the wrangler error message actually lives. - The online specs import
test/expectfromapp/e2e/fixtures/online.tsinstead of@playwright/test. ItsonlineWorkerAlivefixture isauto, pollsGET /api/v1/healthbefore each test, and fails withInfra: … o wrangler dev caiuwhen the Worker is gone — so an infra death is labelled as one instead of being misattributed to the screen under test. New online specs must import from that fixture, not from@playwright/test. retries: 1covers runner contention (cold Vite, timeouts), not a dead Worker: every attempt of a spec whose Worker is gone still fails the health poll. The specs are retry-safe by construction (09 re-picks a free student, 13/14 useuniqueNameand stop before the check-in is effective); check that property before adding a spec that mutates shared state.
- The
Cost of these tiers, and why changing them is harder than it looks
The wall clock of the CI, where it actually goes (measured on the runner), the cost model of the unit tiers (≈1.3 s per jsdom file, of which only ~0.13 s is real test work), and every "obvious" optimization that was measured and rejected — workers: 2, higher maxWorkers, happy-dom, isolate: false, deps.optimizer, trimming specs — are documented in docs/research/ci-cost-model-contention-and-levers-2026.md. Read it before touching the tier configuration or adding a spec to the smoke subset: the short version is that the CI is contention-bound, not work-bound, so parallelism knobs do nothing, and the remaining levers are structural (a second runner, selective runs).
13. Ownership and Updates
Contributors updating requirements or architecture must update this handbook if execution expectations change.
Do not leave process-critical changes undocumented.