Usage
Benchmark guide
copy markdownPoros Bench is an operational harness for comparing Poros blockwise against resident QLoRA, Unsloth, PEFT, and Accelerate device-map offload under a shared configuration (model, LoRA shape, sequence length, steps, batch size, accumulation).
It is not a training-protocol reproducer: the arms share configuration, not a data pipeline (see what is and is not held identical), so treat its s/step cells as indicative. The published comparison tables were produced by dedicated campaign protocols with their own committed artifacts (validation index).
Running a benchmark
# Single model, single card
poros bench run configs/bench/qwen2_5_32b_3090.yaml
# Full architecture x scale matrix
poros bench matrix configs/bench/launch_matrix.yaml
# Compare two runs
poros bench compare ./bench_output/run_a ./bench_output/run_b
# Generate an HTML report from artifacts
poros bench report --html ./bench_outputBenchmark specs
A benchmark spec is a YAML file defining the hardware, models, methods,
training parameters, and reporting options. See configs/bench/ for
examples.
spec_id: qwen2_5_32b_3090
hardware:
gpu_name: "NVIDIA GeForce RTX 3090"
vram_gb: 24.0
pcie_gen: 4
models:
- model_id: "Qwen/Qwen2.5-32B"
display_name: "Qwen2.5-32B"
param_count_b: 32.5
methods:
- poros
- resident
- unsloth
- peft
- accelerate_offload
precisions:
- nf4
training:
steps: 200
seq_len: 512
batch_size: 1
gradient_accumulation_steps: 8
dataset: "synthetic" # deterministic, license-free tokens; or any HF dataset id
seed: 42
lora:
rank: 16
alpha: 32
dropout: 0.0 # paper parity protocol uses dropout=0.0 (isolates the streaming
# mechanism from dropout RNG); the API default is 0.05
poros:
block_size: 4
enable_prefetch: true
reporting:
output_dir: "./bench_output/qwen2_5_32b_3090"
html: true
markdown: true
emit_json: trueFairness guards
What the harness holds identical across methods:
- Same
seq_len,lora_rank,lora_alpha,target_modules,steps,token_count, andmodel_idacross all methods in a run: every arm is generated from one shared spec, so these values are structurally single-valued rather than checked after the fact. - Every method runs in a fresh process (subprocess isolation) to prevent shared-state contamination.
- VRAM is measured via
torch.cuda.max_memory_allocated(), reset between methods. (Per-step reserved-memory samples appear in a training run'sevents.jsonl, not in benchmark cells.)
What it does NOT hold identical -- read before citing s/step numbers:
- Data pipeline. The baseline arms train on a fixed synthetic batch
reused every step; the Poros
trainerentrypoint trains on the spec's real dataset throughPorosTrainer(tokenization, per-step host-to-device transfer). This biases step time against Poros, not for it, but the arms are not step-for-step identical. - Optimizer protocol. Baseline arms use a bare AdamW loop; the Poros trainer arm uses the released pipeline's LR schedule and clipping. The spec carries no shared optimizer/scheduler contract.
VRAM comparisons are insensitive to both differences; s/step comparisons are indicative, not protocol-controlled.
Metrics collected
For each (method, model, precision) cell:
- VRAM: peak allocated (MB), plus per-step averages
- Speed: seconds per step, tokens per second
- Parity: max |delta| on loss, max |delta| on adapter weights, bitwise exact flag
- Status: ok, oom, error, skipped
Output artifacts
Each run produces:
spec.json— the benchmark spec used{method}_{arch}_{precision}.json— oneRunArtifactper cell (the environment snapshot is a field on each artifact, not a separate file)results.jsonl— every cell appended as it completes, so a partial run is still readablereport.html— self-contained HTML report (ifhtml: true)report.md— Markdown report (ifmarkdown: true)
Per-step training events come from the trainer, not this harness: look in
{output_dir}/events.jsonl of a poros train run.
Speed and VRAM cells carry no parity fields, because measuring parity means
running the blockwise and resident arms of one configuration against each
other — which the parity gates do on their own. So an empty parity value
means "not measured here", never "measured as zero", and the report prints
-- rather than 0.00e+00.
Available benchmark configs
| Config | Description |
|---|---|
configs/bench/qwen2_5_7b_3090.yaml | Qwen2.5-7B on RTX 3090 |
configs/bench/qwen2_5_32b_3090.yaml | Qwen2.5-32B on RTX 3090 (headline demo) |
configs/bench/launch_matrix.yaml | Launch matrix (Qwen2.5-32B + Qwen2.5-72B + Qwen3-32B) |
Methodology
The A100, B300, and canonical-campaign numbers come from dedicated campaign
runs whose raw receipts are committed (validation index);
today's poros bench shares their configuration schema but not their
step-for-step protocol (see the fairness guards above). The RTX 5090 rows
use the instrumented single-pass validation protocol, and the RTX 3090
same-card table is the paper's Table 4. The campaign methodology:
- Each method trains under identical LoRA config, sequence length, batch size, gradient accumulation, and seq-length token budget.
- VRAM is the peak
max_memory_allocatedduring training, measured after atorch.cuda.reset_peak_memory_stats()call. - Parity is the elementwise max absolute difference between Poros blockwise and the fully-resident reference, measured on loss scalars and adapter weight tensors at every step.
- Speed is wall-clock seconds per optimizer step: the median of post-warmup step times within a pass, then the median across passes (5-step warmup).
- Unsloth runs in a fresh subprocess to prevent shared cuBLAS/allocator state from contaminating its VRAM measurement.
No number in these tables is derived from ratios, extrapolated, or fabricated: every cell corresponds to a measured run with artifacts. Where the docs do extrapolate — architecture coverage inferred from one size in a family, or a "fits on N GB" column derived from peak VRAM — they say so at the point of the claim.