> For the complete AWA documentation index, see [`llms.txt`](../llms.txt).

# Upgrading from 0.6 to 0.7

The 0.6 → 0.7 upgrade is short: **the storage step is "be finalized."** Everything else in
0.7 is additive.

## Restricted runtime roles and v047

v047 refreshes lane-sequence helpers in the default and existing custom
queue-storage schemas without changing function signatures or cursor semantics.
Existing provisioned lanes run without schema CREATE; single-role owners retain
lazy creation. The migration adds no runtime version gate or authority flip.

A [migrate-first rehearsal with released `awa-pg==0.6.7`](https://github.com/hardbyte/awa/pull/501#compatibility-evidence)
verified a restricted worker completing jobs on the upgraded schema. That
rehearsal does not cover concurrent mixed-version rollouts.

For a runtime without DDL privileges, apply v047 as the migrator, run
`awa storage prepare-queue --queue <name>` for every physical queue, and grant
sequence `USAGE, SELECT, UPDATE` before starting producers or workers. Match
shard and stripe counts and provision extra lanes before increasing them.
See [Database roles](../security/database-roles/index.md#provision-queues-before-starting-restricted-runtimes).

## The one gate

`awa migrate` on a 0.7 binary refuses to apply migrations unless one of these holds:

- the cluster's storage transition is **finalized** (`awa storage status` reports
  `state = active`, `current_engine = queue_storage`), or
- the install is **fresh** — a brand-new database, or a schema with no jobs and no
  recently-live runtimes (the same conditions that let a fresh install auto-finalize at
  worker startup).

Any other shape — canonical with work, `prepared`, `mixed_transition` — is refused with an
error naming the steps below. Nothing is applied on refusal; re-run `awa migrate` after
finalizing.

This is the [ADR-037](../adr/037-canonical-engine-deprecation/index.md) deprecation gate: the
canonical (row-mutating) engine is deprecated in 0.7 and its claim/execution/trigger paths
are removed in 0.8. Requiring finalization at the 0.7 boundary means no 0.7 feature ever
needs a canonical implementation, and gives operators one unambiguous instruction.

## If you are on 0.6, finalized

You are done with the storage step. Deploy 0.7 binaries and run `awa migrate` as usual.

## If you are on 0.6, not yet finalized

Complete the staged transition **on your 0.6 binaries** first
(full procedure: [Upgrade 0.5 to 0.6](../upgrade-0.5-to-0.6/index.md)):

```bash
awa storage prepare --engine queue_storage
awa storage enter-mixed-transition
awa storage finalize --wait
```

For a stopped fleet, a CLI containing #457 can replace the middle command
with `awa storage enter-mixed-transition --quiesced`; follow the
[quiesced procedure](../upgrade-0.5-to-0.6/index.md#quiesced-alternative-to-the-witness-runtime).
If canonical callbacks are outstanding, deploy the 0.6 backport of #462 before
flipping routing. Neither fix requires a new schema migration.

Then upgrade binaries to 0.7 and run `awa migrate`.

If `finalize --wait` sits at a non-zero backlog that never falls while jobs are visibly executing, your workload probably contains **perpetually snoozing jobs** — handlers that end every run in `JobResult::Snooze`. On builds before [#456](https://github.com/hardbyte/awa/issues/456) those re-entered canonical `scheduled_jobs` after each post-flip run, replenishing the backlog forever. Roll to a build carrying that fix (or apply the [manual backlog migration](../upgrade-0.5-to-0.6/index.md#known-issues)) before waiting on finalize.

Remember the transition is a one-way door once queue-storage work is accepted; the 0.5→0.6
guide covers the abort boundaries.

## If you are on 0.5.x

Step through 0.6: upgrade to the latest 0.6.x, run `awa migrate`, walk the staged
transition to `finalize`, then upgrade to 0.7. A 0.7 binary will not migrate a 0.5-era
schema directly — stepping-stone upgrades are the supported path (the same pattern Oban,
River, and Postgres itself use).

## The v043 ring-rotation ledger migration

Migration **v043** ([#371](https://github.com/hardbyte/awa/issues/371),
[ADR-040](../adr/040-append-only-ring-rotation-ledger/index.md)) moves the queue/lease/claim ring
cursors out of the mutable `{ring}_ring_state` singleton columns (`current_slot`,
`generation`) into append-only `{ring}_ring_rotations` ledgers (dead-tuple-free rotation
under a pinned MVCC horizon). It is delivered as a staged **expand → flip → contract**
upgrade supporting a mixed 0.6.2/0.7 fleet.

> **Required stepping-stone:** first roll the whole fleet to **0.6.2 or later**.
> 0.6.2 is the first 0.6 build that recognizes v043 as forward-compatible only
> while ring authority is `columns`, refuses unknown newer schemas without
> modifying them, and refuses to restart after the authority flip. Do not run a
> 0.7 migration while a 0.6.0 or 0.6.1 binary can invoke `awa migrate`; those
> releases contain the destructive newer-schema misclassification fixed by
> [#392](https://github.com/hardbyte/awa/issues/392).

`awa migrate` enforces this stepping-stone: v043 is refused while any runtime with a
fresh heartbeat reports a version below 0.6.2 or an unparseable version. The
`--allow-live-runtimes` override is available for operators who have independently verified
the fleet; it should not be needed in the normal rollout.

**v043 is additive (the expand phase).** It creates and seeds the three ledgers and the
rollup-delta landing table, and it **keeps** the compat `current_slot` / `generation`
columns in place. Each queue-storage schema gets a `ring_cursor_authority` control row
that selects which representation is authoritative:

- `columns` (compat) — the pre-0.7 singleton columns are authoritative, exactly as 0.6
  wrote them, so a live 0.6.2 binary keeps working. 0.7 rotators additionally shadow every
  advance into the ledger, keeping it a ready-to-promote copy.
- `ledger` — the append-only ledgers are authoritative (the #371 dead-tuple win).

An **upgrade** starts in `columns`; a **fresh install** starts directly in `ledger` (no
old binary can exist). The flip is one-way.

### Procedure (either order works after the 0.6.2 stepping-stone)

First roll 0.6.2 across the fleet as a normal patch release. After that, both
orders are supported:

1. **Roll binaries, then migrate**, or **migrate, then roll binaries** — either way the
   fleet runs mixed 0.6.2/0.7 against one database once v043 is applied. In
   compat mode a 0.6.2 rotator and a 0.7 rotator serialize on the same
   `{ring}_ring_state` row lock, so the cursor stays correct. A 0.6.2 restart also
   recognizes v043 while authority remains `columns`; it refuses rather than mutating the
   schema once authority is `ledger`.

   If you roll binaries first, a 0.7 worker **refuses to start** on the
   pre-migration v040 schema — it fails closed at startup naming `awa migrate`
   as the fix, and completes no work. Under a rolling deployment the new pods
   crash-loop harmlessly while the remaining 0.6.2 workers keep draining
   traffic, and come up unaided the moment the migration commits. Plan
   capacity for that window, or migrate first to avoid it.
2. Once **every** worker is on 0.7, promote to ledger authority to unlock the full #371
   dead-tuple benefits — either:
   - manually: `awa storage flip-ring-authority` (add `--schema <name>` for a custom
     schema; `--check` prints the fleet flip-readiness and exits without changing
     anything), or
   - automatically: the maintenance leader auto-flips once every fresh-heartbeat runtime
     has reported a 0.7+ `binary_version` continuously for a stable period (default 10
     minutes; a builder knob). It logs the flip loudly.

If you migrate first, roll 0.7 workers promptly. An all-0.6.2 fleet on v043 remains safe:
crash and heartbeat rescue are batch-aware, but 0.6.2's deadline-rescue sweep does not read
the v042 compact claim batches used for newly claimed deadline jobs. Deadline-based rescue
for those jobs resumes when a 0.7 runtime takes maintenance leadership. Roll the
0.6 maintenance leader promptly after migrating. Awa's built-in
migrator applies v041-v043 in one transaction, so other sessions see v040 or v043, never an
intermediate version. An external migration runner that commits each version separately can
expose v041/v042; a restarting 0.6.2 worker refuses those transient schemas and connects
normally once v043 is committed.

The manual flip **refuses** (without `--force`) while any fresh-heartbeat runtime is not
known to be flip-aware — i.e. a 0.6 (or pre-flip 0.7) binary might still be reading the
compat columns. Roll the whole fleet first, or pass `--force` only once you have confirmed
no pre-flip binary is live.

### After the flip — do not roll back to a pre-flip binary

The flip **fences returning pre-flip binaries**: under the three cursor locks it first
reconciles and verifies every shadow ledger, then poisons both compat cursor fields and
the legacy prune metadata. A database trigger rejects any later old-style cursor advance,
including the `-1 -> 0` rotation a pre-ledger binary would otherwise compute. A 0.6.2 (or
pre-flip 0.7) binary that reconnects therefore fails loudly instead of misrouting writes
or pruning an authoritative ledger slot.
Rolling a binary back **across the flip** is therefore not supported; roll back only to
another 0.7 build (which reads the ledger). **Before** the flip, the normal one-minor
skew guarantee holds and rollback to the 0.6.2 stepping-stone is safe.

### Contract (0.8)

Dropping the compat columns for good is deferred to **0.8** (the contract phase), after
ledger authority is universal and the release's capability gate excludes pre-flip binaries.
Until then the columns remain as a cheap, cold safety net.

Fresh installs are unaffected — they get ledger authority directly and never touch the
compat path.

## Runtime deprecation warning

Any 0.7 worker whose effective storage resolves to the canonical engine — an unfinalized
cluster, an explicit `canonical_drain` role, or a builder configured for canonical storage —
logs a startup warning naming this guide. Treat it as a to-do, not an emergency: canonical
still works in 0.7, and is gone in 0.8.

## Rollback

0.7's migrate *gate* itself changes nothing in the database. Migration **v043** is
additive (it keeps the compat cursor columns), so applying it does **not** break the
one-minor-version skew window: a 0.6.2 binary keeps working against a v043 schema in compat
authority. The skew window only ends at the **ring-authority flip** (see above): after the
flip the stale compat columns are poisoned, so a pre-flip binary fails loudly and rolling
back across the flip is **not** supported — roll back only to another 0.7 build. Before
the flip, rollback to the 0.6.2 stepping-stone is safe. There is no schema downgrade path (unchanged from
previous releases).


## v046: wait-free admin dirty-key marks

v046 replaces the keyed `admin_dirty_queues` / `admin_dirty_kinds` tables the
canonical triggers wrote with `INSERT ... ON CONFLICT DO NOTHING` by append-only
`admin_dirty_queue_marks` / `admin_dirty_kind_marks` tables with no index or
constraint, and rewrites `mark_dirty_keys_*`, `recompute_dirty_admin_metadata()`
and `refresh_admin_metadata()` around them. The trigger write inside every job
transition becomes a plain heap insert that cannot wait on another transaction,
and the full refresh no longer `TRUNCATE`s under `ACCESS EXCLUSIVE`.

No operator action is required. The migration creates two small tables, copies
pending marks, and replaces function bodies; it takes no lock that conflicts with
job traffic. 0.6.x runtimes call the maintenance functions by name and keep
working on the migrated schema in either order. External SQL runners apply the
file like any other additive migration. The old tables are left in place and are
dropped by a later contract migration.

## v045: opt-in periodic ownership

v045 adds ownership, retirement, and normalized desired-declaration control
state. Existing unowned additive schedules retain their behavior. Apply the
migration before starting an authoritative runtime; a binary-first authoritative
startup refuses until the schema is present. No existing schedule is adopted or
retired by migration.

Roll every runtime to a cron-protocol-capable build before expecting automatic
retirement. Fresh older instances conservatively block it even if their cron
configuration is unrelated. Then explicitly adopt the owner's existing schedule
names and enable complete-set registration. Inspect `awa cron plan OWNER` before
removing declarations. A stopped fleet is an outage, never a delete signal;
use explicit `cron retire-owner OWNER --apply` for decommissioning.

Use current operator clients for owned schedules. The released 0.6.7 automatic
enqueue and additive UPSERT paths are fenced on retired rows; old manual trigger
uses a separate read/insert path and is outside the owned-schedule operator
contract. Old physical deletion is rejected. Existing jobs and retries are not
cancelled. Restore uses current database time, without replaying retired time.

The Rust migrator drains cron enqueues before applying the pending range, holding
`cron_jobs` in `ACCESS EXCLUSIVE` mode through commit. This prevents a lock-order
cycle between released cron enqueue and earlier storage DDL. Cron evaluation
pauses for the whole range (several seconds when upgrading from v040), then
resumes under its missed-fire policy. Ordinary workers may still encounter the
existing storage-DDL contention handled by the migration retry policy.

External SQL runners crossing v045 in a transaction with earlier migrations must
likewise acquire `LOCK TABLE awa.cron_jobs IN ACCESS EXCLUSIVE MODE` before the
**first** pending migration, when the table already exists. Acquiring it only at
v045 is too late. A fresh install has no old cron enqueues to drain.

External SQL runners must apply the complete migration transactionally under
the migration lock and retain the runtime-snapshot statement trigger: it makes
old and new evidence writers share the same retirement serializer. Do not disable
it or write declaration state independently of the snapshot transaction. The
lock protects control-plane decisions only, not ordinary job/lease heartbeats.

Reproduce the released-artifact proof with
`scripts/rehearse-cron-ownership.sh` (disposable DATABASE_URL required); the cron
model suite is `scripts/check-cron-models.sh`. Existing ring-authority rollback
restrictions still apply independently of cron ownership.

Reintroducing a retired desired name in a later deploy does not fail startup.
It remains inert and visible as an ownership/retirement blocker until an operator
restores it. Use `awa cron restore NAME` or `awa cron restore-owner OWNER` to
preview, then add `--apply`. Restore-owner affects only that owner's retired
schedules, preserves pause state, and starts evaluation from current database
time. Foreign ownership still fails startup and requires explicit transfer.
During mixed-manifest rollouts, existing definitions retain their last agreed
value; new names can be inserted immediately. Definition updates wait for fresh
capable fleet agreement, while removals additionally wait through grace.
