`table.extract()` built its lookup table with
`INSERT OR IGNORE ... SELECT DISTINCT <cols> FROM <table>`, which included
the all-NULL combination. That created a spurious lookup row for NULL and
pointed every NULL source row at it, instead of leaving those rows with a
NULL foreign key.
Before:
db["creatures"].extract("type")
# type lookup: [{"id": 1, "type": None}, {"id": 2, "type": "dog"}]
# creatures: Simon -> type_id=1, Natalie -> type_id=1, Cleo -> type_id=2
After:
# type lookup: [{"id": 1, "type": "dog"}]
# creatures: Simon -> type_id=None, Natalie -> type_id=None, Cleo -> type_id=1
A row whose extracted columns are entirely NULL represents "no value", so it
now keeps a NULL foreign key and no lookup row is created for it. The fix adds
a `WHERE NOT (<col> IS NULL AND ...)` guard to the lookup INSERT; the existing
`IS`-based foreign-key UPDATE then leaves those rows NULL automatically (the
subquery finds no matching lookup row).
For multi-column extracts, only the fully-NULL combination is skipped — a
partial-NULL combination (some extracted columns set, others NULL) is a
genuine distinct value and is still extracted and shared between matching rows.
Updates test_extract_works_with_null_values to assert the corrected behaviour
and adds regression tests for the single-column and multi-column cases.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Closes#577
This should solve all sorts of problems seen by users of platforms that throw errors on writable_schema.
Also added `add_foreign_keys=` and `foreign_keys=` parameters to `table.transform()`.
Takes my test down from ten minutes to four seconds!
* Removed unnecessary update() optimization
* Added column_order= to .transform() and .transform_sql()
* Tests for reusing lookup table in extract()
Closes#172