Benchmarking Notes¶
This document captures the benchmark harnesses used in the repo and a few reference results from named development and comparison environments.
Reference Environments¶
- Developer laptop reference: Apple M5 MacBook Air (16 GB), PostgreSQL 17 in Docker (OrbStack). Several awa-only regression examples below come from this environment.
- Dedicated enqueue reference: PostgreSQL 17 on a separate Linux host. These runs are useful for producer-path shape comparisons and are not published throughput claims.
- Cross-system comparison reference: postgresql-job-queue-benchmarking. That repo owns fair-comparison hardware, adapter configuration, raw output, and result summaries. Its default Postgres image is 17.2;
--pg-image postgres:18-alpineselects PG18. - Example commands assume a developer database URL:
postgres://postgres:test@localhost:15432/awa_test - Benchmarks live in:
awa/tests/benchmark_test.rsawa/tests/scheduling_benchmark_test.rsawa/tests/failure_benchmark_test.rs - Python worker benchmarks live in:
awa-python/scripts/benchmark_runtime.py - Shared output schema:
awa/tests/bench_output.rs(Rust)awa-python/scripts/bench_output.py(Python)
These are engineering benchmarks, not published vendor-style numbers. The main goal is to compare shapes, validate architecture changes, and catch regressions. Treat each result as tied to the environment named beside it; do not compare numbers from different sections as though they came from one machine.
North-Star Metrics¶
Awa's performance target is sustainable end-to-end queue work, not the largest enqueue-only number. A benchmark table should lead with:
- completed jobs/sec — durable terminal transitions, not just handler returns
- p99 end-to-end latency — producer enqueue to durable completion
- queue depth / backlog growth — whether offered load is being absorbed or stored as future latency
- WAL bytes per completed job — the main Postgres durability budget for the hot path
- transaction commits per completed job — connection and commit pressure
- dead tuples and prune lag — whether churn stays bounded under overlap readers, retries, and terminal bursts
Producer throughput is still measured, especially for INSERT vs COPY and enqueue_shards sweeps, but it is a secondary producer-path diagnostic. A configuration that can enqueue 50k/s and complete 5k/s is a backlog and latency problem unless the workload is intentionally bursty and has enough headroom to drain before the SLA window closes.
Positioning Proof Checklist¶
Queue storage changes Awa's comparison set.
The interesting public question is no longer "can Postgres run jobs?" It is "can a Postgres job queue keep dispatch latency low while also keeping hot-path dead tuples bounded?"
That means the benchmark burden for positioning is different from the benchmark burden for internal regressions. Before making the strongest public claim, the comparison needs to be rerun cleanly on the same hardware against the right reference points:
- Awa queue storage
- PgQue
- River
- optionally Oban Pro as a paid partitioned reference
The benchmark set should include:
- idle pickup latency
- sustained runtime throughput and durable completion throughput
- overlap readers / MVCC horizon pressure
- mixed workload soak
- terminal-failure burst
- WAL bytes/job, commit pressure, queue depth, and dead-tuple pressure for every headline table
Operational differences should be called out, not hidden:
- Awa does not require
pg_cron; the worker runtime owns dispatch, rescue, rotation, and prune. - PgQue is designed around a periodic ticker and recommends
pg_cron. - River and Oban are job frameworks rather than shared-log event queues, so they remain the more comparable references for worker semantics.
Commands and raw output for any public-facing table should live in-repo under docs/adr/bench/ or another stable location so the resulting claim is auditable.
Methodology Notes¶
The most important lesson from this round of work is that benchmark isolation matters.
The maintenance service is global to the worker instance. If a benchmark leaves behind millions of deferred rows in awa.scheduled_jobs, later "hot path" benchmarks may accidentally measure background promotion work as well. The sustained hot/deferred benchmarks reset runtime state first to avoid that pollution.
Two benchmark shapes are used:
- Burst/frontier benchmarks: seed a backlog, then drain it
- Steady/sustained benchmarks: warm up, then measure a fixed time window
Steady numbers are the better indicator of sustained runtime behavior.
MVCC Horizon Benchmark¶
awa/tests/scheduling_benchmark_test.rs also includes an ignored test_mvcc_horizon_overlap_benchmark scenario for the failure mode described in PlanetScale's "Keeping a Postgres queue healthy" post: overlapping long-lived transactions pin the MVCC horizon while queue churn continues.
The benchmark:
- runs a steady producer feeding
awa.jobs_hot - enables aggressive completed-job cleanup to generate update/delete churn
- starts overlapping
REPEATABLE READreader transactions on separate connections to hold old snapshots open - samples per-second throughput, queue depth, and
pg_stat_user_tables(n_dead_tup, vacuum counters) forawa.jobs_hot
Reader modes:
idle_snapshot: pins a snapshot and waits inside the transaction; useful for the pure MVCC-horizon caseactive_scan: runs repeated queue scans inside the open transaction so the reader behaves more like overlapping analytics work on the primary
Example:
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
cargo test --release --package awa --test scheduling_benchmark_test \
test_mvcc_horizon_overlap_benchmark -- --ignored --exact --nocapture
This scenario is also wired into .github/workflows/nightly-chaos.yml and runs on PostgreSQL 18 in the Rust nightly benchmark lane.
The nightly lane uses a shorter CI profile so the benchmark stays cheap on shared runners while still exercising overlap readers and cleanup pressure:
AWA_MVCC_JOB_RATE=400AWA_MVCC_BASELINE_SECS=5AWA_MVCC_OVERLAP_SECS=12AWA_MVCC_COOLDOWN_SECS=5AWA_MVCC_OVERLAP_READERS=2AWA_MVCC_OVERLAP_HOLD_SECS=12AWA_MVCC_OVERLAP_STAGGER_SECS=4AWA_MVCC_READER_MODE=active_scanAWA_MVCC_ANALYTICS_TICK_MS=500
Nightly regression checks use per-run ratios (overlap_handler_per_s and cooldown_handler_per_s relative to the same run's baseline window) plus guardrails on dead_tup_delta and max_available.
That nightly profile is intentionally short. It checks that Awa still behaves reasonably under overlapping analytical readers, but it is not the same thing as the 15-minute mixed-workload soak discussed in the PlanetScale post. For a closer reproduction, increase duration and reader hold times and keep the readers in active_scan mode so queries overlap continuously rather than simulating only idle in transaction.
A second ignored benchmark target now exists for that longer profile: test_mvcc_horizon_planetscale_soak. Its default shape is closer to the blog's mixed-workload setup:
800jobs/sec producer rate3overlapping analytics readers120sreader hold time20sstagger between readers15moverlap window withactive_scan
That soak benchmark is wired into CI as a weekly/manual run, while the shorter test_mvcc_horizon_overlap_benchmark remains the daily nightly smoke.
MVCC bench knobs¶
AWA_MVCC_JOB_RATE— steady producer rate in jobs/secAWA_MVCC_BASELINE_SECS/AWA_MVCC_OVERLAP_SECS/AWA_MVCC_COOLDOWN_SECSAWA_MVCC_OVERLAP_READERS/AWA_MVCC_OVERLAP_HOLD_SECS/AWA_MVCC_OVERLAP_STAGGER_SECSAWA_MVCC_READER_MODE—idle_snapshotoractive_scanAWA_MVCC_ANALYTICS_TICK_MS— cadence for repeated analytics scans inactive_scanmodeAWA_MVCC_CLEANUP_INTERVAL_MS/AWA_MVCC_CLEANUP_BATCH_SIZE
Cross-system comparison¶
The awa benches in this file are the awa-only regression detection track — fast, precise, hard thresholds, awa code only. They are not designed for fair comparison with other systems.
For cross-system comparison, the long-horizon harness lives in a separate repository so each system-under-test can be benchmarked through its public API on equal footing:
hardbyte/postgresql-job-queue-benchmarking
The companion repo currently benchmarks awa against pgque, procrastinate, pg-boss, river, oban, and pgmq using a shared subprocess contract: each adapter is a self-contained binary that emits one JSON sample per line. Headline numbers, per-system architectural notes, and reproducible run instructions are in that repo's SYSTEM_COMPARISONS.md and results/2026-04-28/SUMMARY.md.
The two tracks are deliberately separate. If they ever diverge on workload shape or thresholds, the awa-only benches in this file are canonical for awa's own numbers and the cross-system runner defers.
Python Runtime Benchmarks¶
The Python benchmark script exercises the real awa-python worker path while reusing the same database-facing benchmark shapes as the Rust runtime:
Baseline scenarios (--scenario baseline):
copy: Python clientenqueue_many_copydirect queue-storage COPY throughputhot: sustained worker throughput over pre-seededawa.jobs_hotscheduled: sustained deferred promotion over pre-seededawa.scheduled_jobs
Failure scenarios (--scenario failures):
terminal_1pct/10pct/50pct: terminal failuresretryable_1pct/10pct/50pct: retry-once failurescallback_timeout_10pct: callback registration with short timeoutmixed_50pct: rotating through terminal, retryable, and success modes
The worker-focused scenarios seed with SQL directly so the reported number is about Python handler dispatch and runtime behavior, not enqueue serialization.
Reference Results¶
Immediate Enqueue Throughput¶
The enqueue path that is being measured here has a few important architectural properties:
- homogeneous inserts now route directly to
awa.jobs_hotorawa.scheduled_jobsinstead of going through the compatibility view - admin metadata maintenance moved from row-level triggers to statement-level trigger batches
- COPY staging now reuses a session-local temp table and stages typed values instead of reparsing text on the final
INSERT ... SELECT
Example reference result from the developer laptop reference:
- developer laptop reference (
Apple M5, release build, v0.5.0-alpha.0): insert_only_single: about30k inserts/scopy_single: about43k inserts/sinsert_contention_distinct(4 producers x 3k): about46k inserts/scopy_contention_distinct(4 producers x 3k, chunk 1000): about100k inserts/sinsert_contention_same_queue(4 producers x 3k): about95k inserts/scopy_contention_same_queue(4 producers x 3k, chunk 1000): about67k inserts/s
These are engineering comparisons, not product guarantees. Their main value is showing where the architecture bottlenecks move as the implementation changes.
Sustained Hot Path¶
Measured with test_runtime_sustained_hot_path after resetting runtime state:
- warmup: 2s
- measurement window: 10s
- queue size seeded: 200,000 immediately-available jobs
Example reference result from the developer laptop reference (release mode, v0.5.0-alpha.0):
- handler returns: about
5.6k jobs/s - DB
completedtransitions: about5.6k jobs/s
This benchmark enables the in-memory OpenTelemetry exporter and the production alerting metrics path (queue depth, lag, wait-duration histogram). Back-to-back A/B testing in the same reference environment shows v0.5.0 is ~44% faster than v0.4.1 (5.9k vs 4.1k/s) — the promotion query optimizations more than offset the metrics instrumentation cost.
Python Runtime Baseline¶
Measured with awa-python/scripts/benchmark_runtime.py on the developer laptop reference database:
enqueue_many_copy: about16.2k jobs/s(50,000jobs in3.09s)- sustained hot path:
- handler returns: about
3.2k jobs/s - DB
completedtransitions: about3.1k jobs/s
The copy scenario measures producer serialization and direct queue-storage COPY. The worker-focused scenarios seed with SQL so the runtime number is not dominated by Python-side enqueue serialization.
Deep Backlog Drain¶
test_queue_storage_deep_backlog_drain_benchmark is an ignored regression benchmark for the spike shape where a large ready backlog drains more slowly than the same fleet processes steady offered load. It seeds with direct queue-storage COPY, optionally analyzes ready_entries, then measures claim/complete throughput and claim latency while draining a fixed backlog.
DATABASE_URL_QUEUE_STORAGE=postgres://postgres:test@localhost:15432/awa_test_queue_storage \
AWA_QS_DEEP_BACKLOG_JOBS=1000000 \
AWA_QS_DEEP_BACKLOG_SECONDS=120 \
AWA_QS_DEEP_BACKLOG_CONSUMERS=16 \
AWA_QS_DEEP_BACKLOG_SHARDS=16 \
cargo test --package awa --test queue_storage_benchmark_test \
test_queue_storage_deep_backlog_drain_benchmark -- --exact --ignored --nocapture
Useful knobs:
AWA_QS_DEEP_BACKLOG_JOBS— ready rows to seed (default100000)AWA_QS_DEEP_BACKLOG_COPY_BATCH— direct COPY chunk size (default1000)AWA_QS_DEEP_BACKLOG_CONSUMERS— concurrent claim/complete loops (default8)AWA_QS_DEEP_BACKLOG_CLAIM_BATCH_SIZE— claim batch size (default512)AWA_QS_DEEP_BACKLOG_SHARDS—awa.queue_meta.enqueue_shardsupserted before seeding (default16)AWA_QS_DEEP_BACKLOG_ANALYZE— runANALYZEafter seeding (default1)
Large Deferred Frontier¶
Measured with test_scheduled_steady_10m_due_1k_per_sec:
- total deferred backlog:
10,000,000rows - due rate target:
1,000jobs/s - measurement window: 10s
Isolated 4-thread Tokio runtime result (release mode, v0.5.0-alpha.0):
9,000of10,000due jobs completed within the window- per-second completions:
0, 1000, 1000, 1000, 1000, 1000, 1000, 1000, 1000, 1000— perfectly steady after the first-tick startup delay - pickup lateness:
p50:229 msp95:332 msp99:343 ms- promotion:
354batches, mean4.3 ms, max59 ms - claim latency: mean
5.0 ms
This demonstrates that the hot/deferred split with literal-state promotion queries handles a 10M-row deferred frontier with steady, predictable throughput.
Key optimization (v0.5.0): Promotion queries use literal state values (e.g., WHERE state = 'scheduled') instead of parameterized (WHERE state = $1). This allows the Postgres planner to match the partial index idx_awa_scheduled_jobs_run_at_scheduled at plan time. With a parameterized query, the planner falls back to a full bitmap scan on multi-million-row tables, degrading promotion from ~4ms to ~400ms per batch (100x slower).
Moderate Deferred Frontier — Higher Due Rate¶
Measured with test_scheduled_steady_2m_due_4k_per_sec:
- total deferred backlog:
2,000,000rows - due rate target:
4,000jobs/s - measurement window: 10s
Result (v0.5.0-alpha.0):
- all
40,000due jobs were picked and completed - pickup lateness:
p50:0 ms,p95:0 ms,p99:57 ms - promotion:
214batches, mean10.0 ms, max176 ms - claim latency: mean
12.4 ms
This validates the architecture at a realistic production scale: 2M deferred rows with 4k/s throughput and reliable promotion.
High-Rate Deferred Frontier: 10M at 6k/s¶
Measured with test_scheduled_steady_10m_due_6k_per_sec:
- total deferred backlog:
10,000,000rows - due rate target:
6,000jobs/s - measurement window: 10s
Result:
58,686of60,000target jobs completed within the window (98%)- per-second completions:
2942, 6834, 6907, 4758, 5915, 5662, 7246, 6446, 5305, 6671 - pickup lateness:
p50:0 ms,p95:310 ms,p99:476 ms - promotion:
242batches, mean6.4 ms, max99 ms
This rate was previously a documented scaling limit (only 20-57k of 60k promoted, promotion at 1.7-3.6s per batch). The literal-state promotion fix (v0.5.0) eliminated the bottleneck entirely.
Promotion and completion throughput knobs:
promote_interval(default250 ms): how often promotion runsAWA_COMPLETION_FLUSH_MS(default1 ms): completion batcher flush intervalAWA_COMPLETION_BATCH_SIZE(default512): max rows per completion-batcher flush. Queue-storage short-job completion uses a fused receipt-claim lock, compact terminal insert, and compact claim-closure insert, so the default keeps finalization latency low while still amortising durable completion work under load.AWA_COMPLETION_SHARDS: number of parallel completion flushers. The runtime default is storage-dependent:8for canonical storage. Queue storage starts at1for ordinary runtimes and uses4once the configured runtime worker capacity is at least512. Queue storage starts conservative because the terminal / receipt path already batches heavily and the effective fleet-wide flusher count isprocesses × AWA_COMPLETION_SHARDS. Override only after measuring end-to-end throughput, p99 latency, WAL bytes/job, and dead tuples for the target worker-process topology.
Internal promotion constants are fixed in the runtime: PROMOTE_BATCH_SIZE = 4,096 rows per promotion batch and PROMOTE_MAX_BATCHES_PER_TICK = 32 batches per maintenance tick. They are not environment variables or builder options.
Concurrent Multi-Queue Lifecycle¶
Measured with concurrent_lifecycle_test.rs. Four queues (email:32, payments:16, analytics:64, webhooks:16 workers), 128 total workers, full insert → claim → execute → complete lifecycle.
Queue-count sweep with 128 total workers, pool=50, 20k jobs (release mode):
| Config | Workers/queue | Throughput | Per-worker |
|---|---|---|---|
| 1 queue × 128 | 128 | ~1.9k/s |
~14/s |
| 2 queues × 64 | 64 | ~2.0k/s |
~15/s |
| 4 queues × 32 | 32 | ~350/s |
~2.7/s |
Key finding: 1-queue and 2-queue throughput is essentially identical, but 4 queues drops 5-6x. The cliff between 2 and 4 queues indicates that the per-queue dispatcher overhead (separate claim query, PgListener, semaphore) compounds non-linearly when many small queues share a pool.
For comparison, the hot-path benchmark (200k pre-seeded into jobs_hot via SQL) reaches ~10k/s because it bypasses insert triggers and has a fully warmed dispatch pipeline. The lifecycle benchmarks exercise the complete path including job-state triggers and notification.
Tuning guidelines:
- Size the connection pool to at least
num_queues * 4 + 20. With 4 queues, that's 36+ connections. - Prefer fewer queues with larger worker pools over many small queues. 2 queues × 64 workers performs the same as 1 × 128.
- Use multi-queue for isolation, priority, or rate limiting — not for throughput.
A unified cross-queue claim query (one SQL round-trip claiming across all queues via LATERAL JOIN) could eliminate the per-queue dispatch overhead. Prototyping shows 16ms for 48 jobs across 2 queues — comparable to single-queue performance. This is tracked as a future optimization.
Progress Feature Overhead¶
ADR-014 introduced the user-facing structured progress API and a two-tier heartbeat flush. In the current queue-storage engine, mutable progress lives in {schema}.attempt_state only after an attempt first needs per-attempt mutable state; short jobs that never report progress do not allocate an attempt_state row. Older benchmark notes refer to progress columns on jobs_hot / scheduled_jobs, which was the canonical-storage implementation.
Performance impact was validated:
- Zero overhead when no progress is set. The heartbeat service partitions jobs by pending progress; jobs without mutations use the heartbeat-only path, and
snapshot_pending_progressreturns empty when no generation has been bumped. - Completion remains narrow. Successful completion clears the progress snapshot while closing the attempt. On queue storage this is handled through the attempt-state / terminal-snapshot path rather than rewriting the ready row.
- Sustained hot-path throughput was unchanged when the progress feature was added. Current hot-path throughput (~5.6k/s) reflects the additional v0.5.0 OTel metrics instrumentation, not progress overhead.
Failure-Mode Benchmarks¶
The failure-mode benchmark suite measures throughput, drain time, and recovery behaviour when a configurable percentage of jobs fail, retry, hang, or trigger rescue paths. This answers a question the happy-path benchmarks cannot: how does failure impact healthy-job throughput?
Benchmark matrix¶
| Scenario | Description |
|---|---|
terminal_1pct / 10pct / 50pct |
N% of jobs fail terminally |
retryable_1pct / 10pct / 50pct |
N% of jobs fail once then succeed on retry |
callback_timeout_10pct |
10% register a callback that times out, then succeed on retry |
deadline_hang_10pct |
10% hang until deadline rescue fires, then succeed on retry |
snooze_once_10pct |
10% snooze once, then succeed |
mixed_all_modes |
50% success, 10% each of terminal/retryable/callback/deadline/snooze |
stale_heartbeat_rescue |
All jobs seeded as "running" with stale heartbeat — measures rescue-to-completion time |
Rust harness¶
Tests live in awa/tests/failure_benchmark_test.rs. Each scenario seeds jobs deterministically by mode, starts a Client with aggressive rescue intervals, drains to terminal states, and emits both human-readable output and a JSONL record. The full matrix command covers the 10 failure scenarios; the stale-heartbeat rescue benchmark is a separate test.
Python harness¶
The Python benchmark (awa-python/scripts/benchmark_runtime.py) supports a failure-mode subset via --scenario failures:
terminal_1pct/10pct/50pctretryable_1pct/10pct/50pctcallback_timeout_10pctmixed_50pct
Python also includes heartbeat rescue via --scenario rescue. It still does not include the Rust-only deadline_hang and snooze_once scenarios. The worker returns RetryAfter, WaitForCallback, Cancel, or raises exceptions based on the job's mode field.
Structured output¶
Both Rust and Python benchmarks emit one JSONL record per scenario, prefixed with @@BENCH_JSON@@ for extraction. Schema version 2:
{
"schema_version": 2,
"scenario": "terminal_10pct",
"language": "rust",
"seeded": 5000,
"metrics": {
"throughput": {
"handler_per_s": 4200.0,
"db_finalized_per_s": 4100.0
},
"drain_time_s": 1.22,
"rescue": {
"deadline_rescued": 0,
"callback_timeouts": 0
}
},
"outcomes": {
"completed": 4500,
"failed": 500
}
}
Extract JSONL from mixed stdout: grep '@@BENCH_JSON@@' output.txt | sed 's/^@@BENCH_JSON@@//'
For enqueue-only benchmarks (insert_only_single, copy_single, and the contention matrix), metrics.enqueue_per_s is emitted instead of metrics.throughput. Those records still include "measurement": "enqueue" in metadata, plus optional Postgres-side deltas such as wal_bytes, temp_bytes_delta, and xact_commit_delta.
Interpreting The Results¶
Some practical guidelines:
- Compare like with like. Burst/frontier benchmarks and steady-state benchmarks answer different questions.
- Reset runtime state before sustained measurements if you want to isolate one path. Global background work can distort results.
- Prefer the
db_completed_deltaview when you care about end-to-end queue completion, not just handler return rate. - Treat the numbers here as environment-specific reference points, not portable guarantees. For cross-system claims, use the companion comparison repo and its recorded environment.
How To Run¶
Happy-path benchmarks¶
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
cargo test --package awa --test scheduling_benchmark_test \
test_runtime_sustained_hot_path -- --exact --ignored --nocapture
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
cargo test --package awa --test scheduling_benchmark_test \
test_scheduled_steady_2m_due_4k_per_sec -- --exact --ignored --nocapture
Enqueue contention benchmarks¶
These are the most useful benchmarks when you want to compare single-producer enqueue against multi-producer contention, or compare chunked INSERT with the COPY staging path under concurrent writers.
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
AWA_BENCH_CONTENTION_PRODUCERS=4 \
AWA_BENCH_CONTENTION_JOBS_PER_PRODUCER=3000 \
AWA_BENCH_INSERT_BATCH_SIZE=1000 \
AWA_BENCH_COPY_CHUNK_SIZE=1000 \
cargo test --package awa --test benchmark_test \
test_enqueue_contention_matrix -- --exact --ignored --nocapture
This emits six JSONL records:
insert_singlecopy_singleinsert_contention_distinctcopy_contention_distinctinsert_contention_same_queuecopy_contention_same_queue
The matrix hard-resets Awa runtime tables before each scenario so later cases do not inherit a larger or dirtier jobs table from earlier ones.
By default, AWA_BENCH_COPY_CHUNK_SIZE should match AWA_BENCH_INSERT_BATCH_SIZE if you want the closest apples-to-apples comparison between chunked INSERT and chunked COPY staging. If you want to test the current "one bulk COPY per producer" shape instead, set AWA_BENCH_COPY_CHUNK_SIZE to AWA_BENCH_CONTENTION_JOBS_PER_PRODUCER.
The optional Postgres profile block in metadata.db_profile is meant to make server runs easier to interpret. In particular:
wal_bytesshows how much WAL the scenario generatedtemp_bytes_deltaandtemp_files_deltashow temp-file pressurexact_commit_deltahelps explain why many small commits degrade throughputtup_inserted_deltashows how much table churn the database observed
Failure-mode benchmarks (Rust)¶
# Full matrix (10 failure scenarios)
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
cargo test --package awa --test failure_benchmark_test \
test_failure_bench_full_matrix -- --exact --ignored --nocapture
# Single scenario
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
cargo test --package awa --test failure_benchmark_test \
test_failure_bench_terminal_10pct -- --exact --ignored --nocapture
# Stale heartbeat rescue
DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
cargo test --package awa --test failure_benchmark_test \
test_failure_bench_stale_heartbeat_rescue -- --exact --ignored --nocapture
Python benchmarks¶
cd awa-python
uv run maturin develop
# Baseline scenarios (copy, hot, scheduled)
PYTHONPATH=scripts DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
uv run python scripts/benchmark_runtime.py --scenario baseline
# Failure scenarios
PYTHONPATH=scripts DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
uv run python scripts/benchmark_runtime.py --scenario failures
# Everything
PYTHONPATH=scripts DATABASE_URL=postgres://postgres:test@localhost:15432/awa_test \
uv run python scripts/benchmark_runtime.py --scenario all
Caveats¶
- Each number is tied to the environment stated near the result. The developer commands use a localhost database URL for convenience, but that URL is not a claim about where every result was produced.
- Ignored benchmark tests are not part of the normal unit/integration test pass.
- The current focus is relative behavior and architectural validation, not cross-machine leaderboard comparisons.
Cross-system numbers¶
Cross-system comparison (awa vs pgque / procrastinate / pg-boss / river / oban / pgmq) lives in hardbyte/postgresql-job-queue-benchmarking. That repo is the source of truth for fair-comparison numbers, chaos results across systems, and the long-horizon harness itself.
The reasoning for the split: cross-system benches need to evolve at the pace of the systems being compared (each gets new releases, new public APIs, new adapter caveats), and the awa-only regression track in this file needs to evolve at awa's pace. Keeping them in one repo coupled their schedules and conflated "did awa regress?" with "is awa still faster than X?". They're separate questions; they're now in separate places.