Follow-up to the #186 fix: skipping the all-NULL combination in the lookup
INSERT is not enough on its own. When extract() reuses a lookup table that
already contains an all-NULL row (e.g. one written by an older sqlite-utils
version or created manually), the IS-based foreign-key UPDATE would still
match that row and link all-NULL source rows to it.
The UPDATE now carries a trailing `WHERE NOT (<col> IS NULL AND ...)` so
all-NULL source rows are never assigned a foreign key and keep the NULL that
the freshly-added column starts with, regardless of the lookup table's
existing contents.
Adds test_extract_all_null_stays_null_with_preexisting_null_lookup_row.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`table.extract()` built its lookup table with
`INSERT OR IGNORE ... SELECT DISTINCT <cols> FROM <table>`, which included
the all-NULL combination. That created a spurious lookup row for NULL and
pointed every NULL source row at it, instead of leaving those rows with a
NULL foreign key.
Before:
db["creatures"].extract("type")
# type lookup: [{"id": 1, "type": None}, {"id": 2, "type": "dog"}]
# creatures: Simon -> type_id=1, Natalie -> type_id=1, Cleo -> type_id=2
After:
# type lookup: [{"id": 1, "type": "dog"}]
# creatures: Simon -> type_id=None, Natalie -> type_id=None, Cleo -> type_id=1
A row whose extracted columns are entirely NULL represents "no value", so it
now keeps a NULL foreign key and no lookup row is created for it. The fix adds
a `WHERE NOT (<col> IS NULL AND ...)` guard to the lookup INSERT; the existing
`IS`-based foreign-key UPDATE then leaves those rows NULL automatically (the
subquery finds no matching lookup row).
For multi-column extracts, only the fully-NULL combination is skipped — a
partial-NULL combination (some extracted columns set, others NULL) is a
genuine distinct value and is still extracted and shared between matching rows.
Updates test_extract_works_with_null_values to assert the corrected behaviour
and adds regression tests for the single-column and multi-column cases.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes#577
This should solve all sorts of problems seen by users of platforms that throw errors on writable_schema.
Also added `add_foreign_keys=` and `foreign_keys=` parameters to `table.transform()`.
Takes my test down from ten minutes to four seconds!
* Removed unnecessary update() optimization
* Added column_order= to .transform() and .transform_sql()
* Tests for reusing lookup table in extract()
Closes#172