From ccaca9936690c6153941186be127a6af425930e5 Mon Sep 17 00:00:00 2001 From: "Joshua (D) Drake" Date: Tue, 4 Aug 2026 19:39:54 -0600 Subject: [PATCH 1/2] Refresh the version-adoption document, and rename it out of two majors Renamed PG18_19_OPPORTUNITIES.md to POSTGRESQL_VERSION_ADOPTION.md. The old name pinned the file to two majors and was already wrong in the filename, which is the least fixable place to be wrong. Every reference to the old name is updated, in ROADMAP.md, gaps/29-read-stream-aio.md and generated_columns.sh. A status table at the top, dated, saying what became of each item rather than only that it closed. Three of them now have measurements behind them: 7. partial-path startup costs (19) measured, changes no plan of ours (#397) 8. parallel autovacuum (19) measured, parameter accepted and ignored (#398) 5. REPACK investigated, and the conclusion was wrong (#399) Item 5 is corrected in place rather than deleted, because how it went wrong is the useful part. It reasoned that REPACK dispatches through relation_copy_for_cluster and that pgColumnar implements that callback. The dispatch was right. The callback is registered and is a stub that raises, so the claim was true of the symbol and false of the behaviour. That is why the file now records what was measured and carries dates. Eight items from the 19 release notes that the document never mentioned, including the one that would have broken us: get_relation_info_hook is removed in 19 and replaced by build_simple_rel_hook. We already gate it, and the gate is now recorded with its file and line, because nothing else explains why the #if is there. A watching section for PostgreSQL 20 with no feature list, deliberately. 19 is not released, master has had one commitfest, freeze is around April 2027, and anything in it can be reverted. The one thing worth watching is whether the static IndexAmRoutines change grows a TableAmRoutine counterpart, which would touch our handler. The suggested order is revised: nothing from 19 is urgent, since both items measured this week came back negative for us. --- design/PG18_19_OPPORTUNITIES.md | 123 ----------------- design/POSTGRESQL_VERSION_ADOPTION.md | 187 ++++++++++++++++++++++++++ design/ROADMAP.md | 4 +- design/gaps/29-read-stream-aio.md | 2 +- test/generated_columns.sh | 2 +- 5 files changed, 191 insertions(+), 127 deletions(-) delete mode 100644 design/PG18_19_OPPORTUNITIES.md create mode 100644 design/POSTGRESQL_VERSION_ADOPTION.md diff --git a/design/PG18_19_OPPORTUNITIES.md b/design/PG18_19_OPPORTUNITIES.md deleted file mode 100644 index 4ba3aca..0000000 --- a/design/PG18_19_OPPORTUNITIES.md +++ /dev/null @@ -1,123 +0,0 @@ -# PostgreSQL 18 and 19 opportunities for pgColumnar - -Features in PostgreSQL 17 through 19 that are not in older releases and that -pgColumnar can use. The support matrix is PostgreSQL 13-19, so anything adopted -must be version-gated with `#if PG_VERSION_NUM` and fall back to the current path -on older majors, in the style of `src/columnar_compat.h`. - -Sources: PostgreSQL 18 release notes (postgresql.org/docs/18/release-18.html); -PostgreSQL 19 beta 1/2 announcements; read_stream.h and read_stream.c -(doxygen.postgresql.org); pgsql-hackers "Trying out read streams in pgvector (an -extension)" and "Allow ReadStream to be consumed as raw block numbers". - -> Status (2026-07): item 1 (read stream / AIO) is shipped -> (`columnar.enable_read_stream`, gap 29). Items 2 and 3 are covered by -> `test/generated_columns.sh` and `test/temporal.sh`. Item 5 (REPACK) has been -> investigated; see the note under that section. - -## 1. Read Stream API + asynchronous I/O — highest value - -- **Availability.** The read stream API (`storage/read_stream.h`, - `read_stream_begin_relation`, `read_stream_next_buffer`) landed in PostgreSQL 17 - and drives sequential scans and ANALYZE. PostgreSQL 18 added the asynchronous - I/O subsystem (`io_method` = sync | worker | io_uring, `io_combine_limit`, - `io_max_combine_limit`, `pg_aios`), and the read stream is the interface that - feeds it. Reported up to 3x on reads from storage. -- **Current state in pgColumnar.** The reader fetches each chunk's value and - exists stream blocks with individual `ReadBuffer` calls, synchronously, one - chunk group at a time (`src/columnar_reader.c`). There is no prefetch, so a - cold scan waits on each block in turn. -- **Opportunity.** Drive the block reads for a stripe/chunk-group scan through a - read stream. The block numbers a columnar scan needs are known ahead of time - from the stripe/chunk catalog, which is exactly the case the read stream (and - `read_stream_next_block` for callers that compute their own block numbers) is - built for. On PostgreSQL 18 this gets AIO prefetch for free; on 17 it gets - posix_fadvise-based prefetch; on 13-16 it falls back to the current path. -- **Effort / risk.** Medium. The API shape has moved between 17, 18, and 19, so - the adoption must be behind a compat shim and validated on each major. Risk is - confined to the read path and covered by the differential/recovery suites. -- **Why first.** It is the largest cold-scan performance lever available and maps - directly onto how pgColumnar already plans its block reads. - -## 2. Virtual generated columns (PostgreSQL 18, now the default) - -- Generated columns can be virtual and are virtual by default; their values are - computed at read time rather than stored. -- **Relevance.** A columnar table with a virtual generated column must return the - computed value on read and must not store a chunk for it. Confirm the table AM - and custom scan handle read-time generation correctly, and add differential - coverage (columnar vs heap) for stored and virtual generated columns on - PostgreSQL 18+. Likely handled at the executor level, but it is unverified. -- **Effort.** Small (a correctness check plus a test), version-gated to 18+. -- **Verified (`test/generated_columns.sh`).** Stored and virtual generated - columns both read correctly on a columnar table across the matrix; the executor - recomputes the virtual value on read, so values match the heap oracle and the - generation expression. One finding: pgColumnar currently *materializes an - all-null chunk* for a virtual generated column at insert rather than skipping - its storage. The read-time value overrides it, so this is a storage - inefficiency, not a correctness problem. Skipping the write for - `attgenerated = 'v'` columns (and returning NULL for them from the reader) is a - worthwhile future write-path optimization; it needs matching reader/vacuum - changes and its own coverage, so it is not bundled here. - -## 3. Temporal constraints (PostgreSQL 18 `WITHOUT OVERLAPS`, PostgreSQL 19 `FOR PORTION OF`) - -- PostgreSQL 18 allows non-overlapping PRIMARY KEY/UNIQUE (`WITHOUT OVERLAPS`) and - temporal foreign keys (`PERIOD`); PostgreSQL 19 adds `FOR PORTION OF` updates. -- **Relevance.** These run through the index and constraint machinery pgColumnar - already integrates with. Verify enforcement on a columnar table and add - coverage; do not assume it works. Effort small, version-gated. - -## 4. btree skip scan (PostgreSQL 18) - -- The btree AM can skip leading index columns. Indexes built on columnar tables - benefit automatically; no pgColumnar code change. Worth a benchmark line and a - test that a multicolumn index on a columnar table plans a skip scan on 18+. - -## 5. REPACK, concurrent (PostgreSQL 19) — investigate - -- PostgreSQL 19 adds `REPACK` for concurrent table repacking, a lower-lock - alternative to `VACUUM FULL`/`CLUSTER`. -- **Relevance.** pgColumnar's `columnar.vacuum`/`vacuum_sorted` rewrite takes - `AccessExclusiveLock`. If `REPACK` dispatches through the table AM (or offers an - extension hook), pgColumnar could offer concurrent compaction. Whether it is - table-AM-extensible is unknown; the first task is to read the REPACK code/tableam - wiring in the 19 tree and decide feasibility. Do not promise it until confirmed. -- **Investigated (PostgreSQL 19 headers).** REPACK is not a new table-AM callback. - `commands/repack.h` shows it reuses the CLUSTER machinery (`cluster_rel`, - `make_new_heap`, `finish_heap_swap`), which dispatches a table rewrite through - the existing `relation_copy_for_cluster` table-AM callback. pgColumnar already - implements that callback (`columnar_relation_copy_for_cluster`), so the - non-concurrent `REPACK` (its default, under AccessExclusiveLock) runs through - the same path as `CLUSTER`/`VACUUM FULL` and should work; this is worth a - direct test. The concurrent variant (`CLUOPT_CONCURRENT`, - ShareUpdateExclusiveLock) captures concurrent changes with logical-decoding - workers (`commands/repack_internal.h`) rather than through a table-AM entry - point, so it depends on logical decoding of the relation's changes and is not - something the AM opts into. Concurrent REPACK on a columnar table is therefore - unverified and likely needs additional work; treat it as future work. - -## 6. Optimizer statistics injection (PostgreSQL 18) - -- `pg_restore_relation_stats()`, `pg_restore_attribute_stats()`, - `pg_clear_relation_stats()`, `pg_clear_attribute_stats()` let code set per- - relation and per-column stats. pgColumnar could seed planner statistics that - reflect columnar reality (per-column distinct/min/max already in the catalog), - improving plan choice. Optional, additive, version-gated to 18+. - -## Noted, out of scope for pgColumnar - -SQL/PGQ property graphs, `ON CONFLICT DO SELECT`, `GROUP BY ALL`, `IGNORE NULLS`, -native JSON `COPY TO`, LZ4 as the default TOAST codec, `pg_plan_advice`. These are -server or SQL-surface features that do not change what a columnar table AM does. -`COPY TO ... (FORMAT json)` (PostgreSQL 19) is unrelated to the Arrow/Parquet -export work in gap 27. - -## Suggested order - -1. Read stream / AIO adoption in the scan (item 1) — flagship performance work. -2. Correctness coverage for virtual generated columns and temporal constraints - (items 2, 3) — small, closes PG18/19 gaps. -3. Investigate REPACK feasibility (item 5). -4. Statistics injection and a skip-scan benchmark line (items 6, 4) as smaller - follow-ups. diff --git a/design/POSTGRESQL_VERSION_ADOPTION.md b/design/POSTGRESQL_VERSION_ADOPTION.md new file mode 100644 index 0000000..0861a89 --- /dev/null +++ b/design/POSTGRESQL_VERSION_ADOPTION.md @@ -0,0 +1,187 @@ +# PostgreSQL version adoption for pgColumnar + +Features in newer PostgreSQL majors that pgColumnar can use, and what has come of +each. Renamed from `PG18_19_OPPORTUNITIES.md` on 2026-08-05: the old name pinned +the file to two majors and was already wrong in the filename, which is the least +fixable place to be wrong (#390, #395). + +The support matrix is PostgreSQL 15 through 19, so anything adopted must be +version-gated with `#if PG_VERSION_NUM` and fall back on older majors, in the +style of `src/columnar_compat.h`. + +Sources: PostgreSQL 18 and 19 release notes; read_stream.h and read_stream.c; +pgsql-hackers threads on read streams in extensions. Every status note carries the +date it was written. A note without one has not been checked. + +## Status, 2026-08-05 + +| item | state | +| --- | --- | +| 1. Read stream and AIO | **shipped**, `pgcolumnar.enable_read_stream` (gap 29) | +| 2. Virtual generated columns | covered by `test/generated_columns.sh` | +| 3. Temporal constraints | covered by `test/temporal.sh` | +| 4. btree skip scan | open, no measurement | +| 5. REPACK | **investigated and the conclusion was wrong**, see below (#399) | +| 6. Statistics injection | open, no measurement | +| 7. Partial-path startup costs (19) | **measured, changes nothing for us** (#397) | +| 8. Parallel autovacuum (19) | **measured, parameter accepted and ignored** (#398) | + +## 19 items this document did not previously mention + +Added 2026-08-05 from the 19 release notes. + +- **`get_relation_info_hook` is removed and replaced by `build_simple_rel_hook`.** + We use that hook, so 19 would have broken the build. Already gated, at + `src/columnar_tableam.c:2443`. Recorded here because nothing else will explain + why the `#if PG_VERSION_NUM >= 190000` is there. +- **Partial-path startup costs are now considered.** Measured against #369's shapes + on 18 and 19: plan selection is identical, so this does not move the grouped + parallel arm. #397 has the numbers. #369 is not version-gated. +- **Parallel autovacuum**, `autovacuum_max_parallel_workers` and the per-table + `autovacuum_parallel_workers`. Measured: a columnar table **accepts and stores** + the storage parameter and our vacuum ignores it, while heap launches workers on + the same fixture. #398. +- **Aggregate processing before joins.** Untested. Interacts with the vectorized + aggregate and its upper-path hook. +- **New planner hooks**: `planner_setup_hook`, `planner_shutdown_hook`, + `joinrel_setup_hook`, `join_path_setup_hook`. We install `set_rel_pathlist_hook` + and `create_upper_paths_hook` today. Untested. +- **Query scans can mark pages all-visible in the visibility map.** We keep our own + visibility map in `columnar_visibilitymap.c`. Untested. +- **`EXPLAIN (ANALYZE, IO)`.** Methodology rather than a feature, and our benchmark + arguments keep coming down to what was actually read. +- **C11 is the minimum language version.** We already set a newer standard, so this + should be free. Worth a preflight rather than an assumption. + +## Watching, PostgreSQL 20 + +**There is no feature list here on purpose.** PostgreSQL 19 is not released: beta 2 +shipped 2026-07-16 and release is planned for September 2026. Master is the 20 +branch with one commitfest run, freeze around April 2027, and anything committed can +still be reverted. A list would be invention. + +One thing is worth watching. 19 changed index access method handlers to a static +`IndexAmRoutines` structure. If a `TableAmRoutine` counterpart follows, it touches +our handler directly. + +## 1. Read Stream API + asynchronous I/O — highest value + +- **Availability.** The read stream API (`storage/read_stream.h`, + `read_stream_begin_relation`, `read_stream_next_buffer`) landed in PostgreSQL 17 + and drives sequential scans and ANALYZE. PostgreSQL 18 added the asynchronous + I/O subsystem (`io_method` = sync | worker | io_uring, `io_combine_limit`, + `io_max_combine_limit`, `pg_aios`), and the read stream is the interface that + feeds it. Reported up to 3x on reads from storage. +- **Current state in pgColumnar.** The reader fetches each chunk's value and + exists stream blocks with individual `ReadBuffer` calls, synchronously, one + chunk group at a time (`src/columnar_reader.c`). There is no prefetch, so a + cold scan waits on each block in turn. +- **Opportunity.** Drive the block reads for a stripe/chunk-group scan through a + read stream. The block numbers a columnar scan needs are known ahead of time + from the stripe/chunk catalog, which is exactly the case the read stream (and + `read_stream_next_block` for callers that compute their own block numbers) is + built for. On PostgreSQL 18 this gets AIO prefetch for free; on 17 it gets + posix_fadvise-based prefetch; on 13-16 it falls back to the current path. +- **Effort / risk.** Medium. The API shape has moved between 17, 18, and 19, so + the adoption must be behind a compat shim and validated on each major. Risk is + confined to the read path and covered by the differential/recovery suites. +- **Why first.** It is the largest cold-scan performance lever available and maps + directly onto how pgColumnar already plans its block reads. + +## 2. Virtual generated columns (PostgreSQL 18, now the default) + +- Generated columns can be virtual and are virtual by default; their values are + computed at read time rather than stored. +- **Relevance.** A columnar table with a virtual generated column must return the + computed value on read and must not store a chunk for it. Confirm the table AM + and custom scan handle read-time generation correctly, and add differential + coverage (columnar vs heap) for stored and virtual generated columns on + PostgreSQL 18+. Likely handled at the executor level, but it is unverified. +- **Effort.** Small (a correctness check plus a test), version-gated to 18+. +- **Verified (`test/generated_columns.sh`).** Stored and virtual generated + columns both read correctly on a columnar table across the matrix; the executor + recomputes the virtual value on read, so values match the heap oracle and the + generation expression. One finding: pgColumnar currently *materializes an + all-null chunk* for a virtual generated column at insert rather than skipping + its storage. The read-time value overrides it, so this is a storage + inefficiency, not a correctness problem. Skipping the write for + `attgenerated = 'v'` columns (and returning NULL for them from the reader) is a + worthwhile future write-path optimization; it needs matching reader/vacuum + changes and its own coverage, so it is not bundled here. + +## 3. Temporal constraints (PostgreSQL 18 `WITHOUT OVERLAPS`, PostgreSQL 19 `FOR PORTION OF`) + +- PostgreSQL 18 allows non-overlapping PRIMARY KEY/UNIQUE (`WITHOUT OVERLAPS`) and + temporal foreign keys (`PERIOD`); PostgreSQL 19 adds `FOR PORTION OF` updates. +- **Relevance.** These run through the index and constraint machinery pgColumnar + already integrates with. Verify enforcement on a columnar table and add + coverage; do not assume it works. Effort small, version-gated. + +## 4. btree skip scan (PostgreSQL 18) + +- The btree AM can skip leading index columns. Indexes built on columnar tables + benefit automatically; no pgColumnar code change. Worth a benchmark line and a + test that a multicolumn index on a columnar table plans a skip scan on 18+. + +## 5. REPACK, concurrent (PostgreSQL 19) -- CONCLUSION WAS WRONG + +**Corrected 2026-08-05 (#399). REPACK does not work on a columnar table.** + +This section previously reasoned that because `REPACK` reuses the CLUSTER machinery, +which dispatches through the `relation_copy_for_cluster` table-AM callback, and +because pgColumnar registers that callback, non-concurrent `REPACK` "should work". + +The dispatch reasoning was right. The conclusion was not, because the callback is +**registered and is a stub**: + +```c +pgcolumnar_relation_copy_for_cluster(...) +{ + COLUMNAR_UNSUPPORTED("CLUSTER / VACUUM FULL"); +} +``` + +So "pgColumnar already implements that callback" was true of the symbol and false of +the behaviour. Measured on 19beta2: `REPACK`, `REPACK ... USING INDEX`, +`REPACK (VERBOSE)`, `CLUSTER` and `VACUUM FULL` all raise, while `REPACK` succeeds on +a heap table on the same build. `REPACK (CONCURRENTLY)` on a table with a primary key +also succeeds on heap and raises on columnar. + +The supported route is `pgcolumnar.vacuum()` and `pgcolumnar.vacuum_sorted()`. + +**This entry is the reason the file now carries dates and states what was measured +rather than what was inferred.** Inspection of which callback a command dispatches +through cannot tell you whether that callback does anything. Pinned by +`test/native_repack.sh`. + + +## 6. Optimizer statistics injection (PostgreSQL 18) + +- `pg_restore_relation_stats()`, `pg_restore_attribute_stats()`, + `pg_clear_relation_stats()`, `pg_clear_attribute_stats()` let code set per- + relation and per-column stats. pgColumnar could seed planner statistics that + reflect columnar reality (per-column distinct/min/max already in the catalog), + improving plan choice. Optional, additive, version-gated to 18+. + +## Noted, out of scope for pgColumnar + +SQL/PGQ property graphs, `ON CONFLICT DO SELECT`, `GROUP BY ALL`, `IGNORE NULLS`, +native JSON `COPY TO`, LZ4 as the default TOAST codec, `pg_plan_advice`. These are +server or SQL-surface features that do not change what a columnar table AM does. +`COPY TO ... (FORMAT json)` (PostgreSQL 19) is unrelated to the Arrow/Parquet +export work in gap 27. + +## Suggested order + +Revised 2026-08-05, after #397, #398 and #399 reported. + +1. **Nothing from 19 is urgent.** The two items measured this week both came back + negative for us: partial-path startup costs change no plan we care about (#397), + and parallel autovacuum accepts a parameter we ignore (#398). +2. **#398 is the only one with a user-visible defect attached**, and the fix is to + reject the storage parameter rather than to honour it. +3. Statistics injection (item 6) and a skip-scan benchmark line (item 4) remain the + open ones, both untested and neither blocking. + +Anything not listed above is a watch item, not planned work. `design/ROADMAP.md` is +the place for planned work, and `docs/roadmap.md` is where a user should be sent. diff --git a/design/ROADMAP.md b/design/ROADMAP.md index 05d72c1..63be49f 100644 --- a/design/ROADMAP.md +++ b/design/ROADMAP.md @@ -21,7 +21,7 @@ matrix. Gap specifications are in [gaps/](gaps/). | Corrupt-input decode/reader hardening | SECURITY_AUDIT.md | | Arrow/Parquet export type coverage (date/time/timestamp/uuid/numeric/json) | gaps/27-IMPL-export-type-coverage.md | | Arrow IPC import (`pgcolumnar.import_arrow`) | gap 27 | -| PG18/19 coverage: generated columns, temporal constraints; REPACK investigated | PG18_19_OPPORTUNITIES.md | +| PG18/19 coverage: generated columns, temporal constraints. REPACK investigated, and the conclusion was wrong: it does not work (#399) | POSTGRESQL_VERSION_ADOPTION.md | | Full index-only scan (visibility-map fork, lazy vacuum, default on) | gap 28 | | Multiple projections (C-Store): catalog, write fan-out, planner scan, vacuum, back-fill | gap 26 piece 2 | | Arrow/Parquet nested export (arrays → List, composite → Struct/group) | gap 27 | @@ -241,7 +241,7 @@ are directions to investigate and spec, not validated recommendations: Features new in PostgreSQL 17-19 that pgColumnar can use, all version-gated to preserve the 15-19 matrix. Detail and sources in -[PG18_19_OPPORTUNITIES.md](PG18_19_OPPORTUNITIES.md): +[POSTGRESQL_VERSION_ADOPTION.md](POSTGRESQL_VERSION_ADOPTION.md): - Read stream / AIO in the scan — shipped, see the Done table. - Virtual generated columns (PostgreSQL 18): confirm read-time generation on a diff --git a/design/gaps/29-read-stream-aio.md b/design/gaps/29-read-stream-aio.md index 7cfd216..4d58901 100644 --- a/design/gaps/29-read-stream-aio.md +++ b/design/gaps/29-read-stream-aio.md @@ -6,7 +6,7 @@ Adopt the PostgreSQL read stream API (PostgreSQL 17+) for the columnar block reads so the PostgreSQL 18 asynchronous I/O subsystem prefetches chunk blocks during a scan. Falls back to the current synchronous path on PostgreSQL 13-16. No on-disk or SQL-surface change; results are unchanged. See -[../PG18_19_OPPORTUNITIES.md](../PG18_19_OPPORTUNITIES.md) item 1. +[../POSTGRESQL_VERSION_ADOPTION.md](../POSTGRESQL_VERSION_ADOPTION.md) item 1. ## Current state diff --git a/test/generated_columns.sh b/test/generated_columns.sh index 37842a5..2b50ecd 100644 --- a/test/generated_columns.sh +++ b/test/generated_columns.sh @@ -80,7 +80,7 @@ if [ "$major" -ge 18 ]; then mism="$(q "SELECT count(*) FROM gv_col WHERE b <> a * 2 OR c <> 'r' || a;")" check "virtual generated: recomputed correctly" "$mism" "0" - # Write-path optimization (PG18_19_OPPORTUNITIES.md item 2): a virtual generated + # Write-path optimization (POSTGRESQL_VERSION_ADOPTION.md item 2): a virtual generated # column is computed on read and never stored, so pgColumnar writes NO column # chunk for it (b, c = column_index 1, 2), while the base column a (index 0) is # stored. The reader returns the column's missing value (NULL) for the absent From cc87e1085f3e9e1f8deefa7db0f900d28c505b80 Mon Sep 17 00:00:00 2001 From: "Joshua (D) Drake" <136637981+ChronicallyJD@users.noreply.github.com> Date: Wed, 5 Aug 2026 08:32:50 -0600 Subject: [PATCH 2/2] docs: item 2 proposed as future work a thing that shipped two weeks ago (#390) The Verified bullet for virtual generated columns said pgColumnar materializes an all-null chunk for them and called skipping that write a worthwhile future optimization. It shipped on 2026-07-22 in "Skip storage for virtual generated columns": columnar_write_state.c skips attgenerated == 'v' outright and columnar_vacuum.c matches. The evidence was already in the file this document cites. generated_columns.sh asserts no chunk is stored for the virtual columns and that the base column still is, three lines below the comment naming this item. So the document proposed as future work a thing done, citing as its source a test that asserts the opposite. That is exactly the failure the rename is for. A status note true when written, carried forward under a filename claiming to be current, is how it gets another year. Every claim in items 2 and 3 is now dated. Item 3 never got a Verified bullet when item 2 did, so the status table said covered while the body said go and verify it and do not assume it works. The table was right; test/temporal.sh exists and runs in the matrix. Item 1 said the read stream falls back on 13-16. The matrix is 15 through 19. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01UqprqkCXuH8SegiZejE1Tw --- design/POSTGRESQL_VERSION_ADOPTION.md | 45 ++++++++++++++++++--------- 1 file changed, 30 insertions(+), 15 deletions(-) diff --git a/design/POSTGRESQL_VERSION_ADOPTION.md b/design/POSTGRESQL_VERSION_ADOPTION.md index 0861a89..9b57e55 100644 --- a/design/POSTGRESQL_VERSION_ADOPTION.md +++ b/design/POSTGRESQL_VERSION_ADOPTION.md @@ -18,8 +18,8 @@ date it was written. A note without one has not been checked. | item | state | | --- | --- | | 1. Read stream and AIO | **shipped**, `pgcolumnar.enable_read_stream` (gap 29) | -| 2. Virtual generated columns | covered by `test/generated_columns.sh` | -| 3. Temporal constraints | covered by `test/temporal.sh` | +| 2. Virtual generated columns | **done**, read and storage both, `test/generated_columns.sh` | +| 3. Temporal constraints | **done**, `test/temporal.sh` | | 4. btree skip scan | open, no measurement | | 5. REPACK | **investigated and the conclusion was wrong**, see below (#399) | | 6. Statistics injection | open, no measurement | @@ -81,7 +81,9 @@ our handler directly. from the stripe/chunk catalog, which is exactly the case the read stream (and `read_stream_next_block` for callers that compute their own block numbers) is built for. On PostgreSQL 18 this gets AIO prefetch for free; on 17 it gets - posix_fadvise-based prefetch; on 13-16 it falls back to the current path. + posix_fadvise-based prefetch; on 15-16 it falls back to the current path. (The + matrix is 15 through 19. An earlier revision said 13-16, from before 13 and 14 + were dropped.) - **Effort / risk.** Medium. The API shape has moved between 17, 18, and 19, so the adoption must be behind a compat shim and validated on each major. Risk is confined to the read path and covered by the differential/recovery suites. @@ -98,24 +100,37 @@ our handler directly. coverage (columnar vs heap) for stored and virtual generated columns on PostgreSQL 18+. Likely handled at the executor level, but it is unverified. - **Effort.** Small (a correctness check plus a test), version-gated to 18+. -- **Verified (`test/generated_columns.sh`).** Stored and virtual generated - columns both read correctly on a columnar table across the matrix; the executor - recomputes the virtual value on read, so values match the heap oracle and the - generation expression. One finding: pgColumnar currently *materializes an - all-null chunk* for a virtual generated column at insert rather than skipping - its storage. The read-time value overrides it, so this is a storage - inefficiency, not a correctness problem. Skipping the write for - `attgenerated = 'v'` columns (and returning NULL for them from the reader) is a - worthwhile future write-path optimization; it needs matching reader/vacuum - changes and its own coverage, so it is not bundled here. +- **Done, and verified by `test/generated_columns.sh`** (status as of 2026-08-05). + Stored and virtual generated columns both read correctly on a columnar table + across the matrix. The executor recomputes the virtual value on read, so values + match the heap oracle and match the generation expression applied to the base + column. +- The storage half is done as well. `columnar_write_state.c` skips + `attgenerated == 'v'` and writes no chunk at all, `columnar_vacuum.c` matches, + and the reader returns the column's missing value (NULL) for the absent chunk, + which the executor then overwrites. Shipped 2026-07-22 in "Skip storage for + virtual generated columns". The suite pins both directions: no chunk for the + virtual columns, a chunk still present for the base column. +- An earlier revision of this document listed that skip as a *worthwhile future + write-path optimization*. It had already shipped, and the suite this document + cites as its evidence asserts the opposite of the claim, three lines from the + comment naming this item. That is the same failure the rename is for: a status + note that was true when written, carried forward under a name that says it is + current. Hence the date above. ## 3. Temporal constraints (PostgreSQL 18 `WITHOUT OVERLAPS`, PostgreSQL 19 `FOR PORTION OF`) - PostgreSQL 18 allows non-overlapping PRIMARY KEY/UNIQUE (`WITHOUT OVERLAPS`) and temporal foreign keys (`PERIOD`); PostgreSQL 19 adds `FOR PORTION OF` updates. - **Relevance.** These run through the index and constraint machinery pgColumnar - already integrates with. Verify enforcement on a columnar table and add - coverage; do not assume it works. Effort small, version-gated. + already integrates with. +- **Done, and verified by `test/temporal.sh`** (status as of 2026-08-05). A + `WITHOUT OVERLAPS` primary key on a columnar table holds the same contents as + its heap twin and rejects the same overlapping rows. `FOR PORTION OF` produces + the same result set, gated to PostgreSQL 19 and visibly skipped below it. +- This item did not get a Verified bullet when item 2 did, so the status table + above said covered while the body still said go and verify it. The table was + right. The gap was in this file, not in the coverage. ## 4. btree skip scan (PostgreSQL 18)