Author
Abstract
Learned cardinality estimation has emerged as a promising alternative to the histogram-based methods embedded in modern query optimizers. Across the last five years, a wide range of query-driven, data-driven, and hybrid estimators have been proposed, each reporting sizeable accuracy improvements on specific benchmarks. The conditions under which these improvements translate into faster query execution, and the cost that they impose on planning and maintenance, remain less well characterized. This paper presents a cross-cutting comparative reanalysis of representative cardinality estimators, drawing performance numbers from the primary literature and four multi-method benchmark studies conducted under comparable PostgreSQL-based evaluation settings. Eight estimators are compared on accuracy; six of the eight also appear in the planning-time and training-cost comparison, for which uniformly measured numbers are available. FACE and FactorJoin are discussed as qualitative reference points, but are not in the quantitative tables. The quantitative analysis centers on JOB-light and STATS-CEB, for which comparable multi-method numbers are publicly available, with the Join Order Benchmark and the TPC-DS / DSB family providing complementary context. We compare estimators along three axes: estimation accuracy, measured by q-error percentiles; planning time and training overhead; and end-to-end query execution time on PostgreSQL with injected cardinalities. Our reanalysis confirms that data-driven estimators achieve the lowest q-error on static workloads but incur planning-time overheads that erode a portion of their end-to-end gains, and that no single estimator dominates across accuracy, latency, and resilience to data updates. Building on this analysis, we propose a workload-driven decision framework that combines a three-axis workload characterisation along query complexity, data scale, and distribution stability, a region-based selection rule, and a break-even cost-benefit model that determines when learned estimators justify their training, planning, and maintenance overhead. We discuss practical implications for analytical data pipelines and identify workload drift and planning-time efficiency as the most pressing open problems.
Suggested Citation
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:dba:jsisia:v:2:y:2026:i:4:p:37-51. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Joseph Clark (email available below). General contact details of provider: https://pinnaclepubs.com/index.php/JSISI .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.