Field values come from vendor documentation, pricing pages and public changelogs, each cell carrying its own provenance. Prices are the list rate for roughly 2 vCPU and 4 GB with no committed spend, normalised to an hourly figure. Several vendors bill per second or per credit and never publish an hourly number, so those cells are marked inferred with the arithmetic noted, or left null where a guess would be worse than saying nothing.
Cold start is the reason this page exists, so read the two columns together. Claimed cold start is whatever the vendor puts on its own homepage, and it is scored at weight zero on purpose — it is a marketing artefact, not a measurement. Vendors quote wildly different things under the same word: time to restore a memory snapshot on already-provisioned hardware, time for the guest kernel to boot, time for the API to return a sandbox ID, or p50 across a pool deliberately kept warm. Measured cold start is what toolweight publishes from its own harness: a single client in a fixed region, TLS handshake included, calling create-sandbox and blocking until a trivial command returns output — the full round trip an agent actually pays for. It runs against an account with no sandboxes alive, after a deliberate idle period, so nothing is pool-warm, and reports p50 and p95 over at least 50 runs so one good afternoon on a vendor's cluster cannot flatter the number. Until that harness has run against every provider here, cells in that column are estimates flagged as inferred, and no cell on this page currently carries measured confidence. We would rather show an honest estimate than launder a vendor's number into a benchmark.
Four rows come from a different kind of source and are flagged accordingly. The capability cells for box, ascii, exe.dev and Islo are transcribed from the comparison table box publishes on its own site, at box.ascii.dev/compare. That is the least neutral source on this page: a chart drawn by a vendor, with that vendor's two products in the first two columns and its competitors arranged around them. toolweight has verified none of it. Every field the table does not address — price per hour, cold start, isolation technology, persistence, regions, SDKs, funding, stars — is left null and unknown rather than filled in by inference.
Three rules keep that table from being laundered into evidence it is not. First, only a cell the table states directly is marked vendor-claimed; where a value had to be derived — reading root access off a "Docker inside the VM" row, or a preview-URL rating off two rows about IP addresses — the cell is marked inferred, because a sound derivation from a vendor's claim is still a derivation. Second, a marketing row is not a measurement: "runs 24/7" is not a published runtime ceiling and "1000+ concurrent VMs ergonomically" is not a published concurrency limit, so those cells are unknown rather than awarded this page's maximum. Third, and most important, absence from a row is scored asymmetrically on purpose. A vendor leaving its own product out of a row is a concession against interest and is credited — box and ascii are both absent from the process-fork and sub-500 ms rows, and that is recorded. A vendor leaving a rival out of a row is the weakest class of claim on this site, so it is recorded as nothing at all: exe.dev and Islo are missing from the snapshot rows, and their snapshot-and-fork cells stay unknown rather than becoming a scored zero on a competitor's say-so. These four rows stay in this shape until they can be re-sourced from each vendor's own documentation, at which point the provenance on every cell changes with them.