Compare commits

...

12 commits

Author SHA1 Message Date
Simon Willison
60740a498b Release 3.21
Refs #348, #364, #366, #368, #371, #372, #374, #375, #376, #379

Closes #380
2022-01-10 18:33:48 -08:00
Simon Willison
f4ea0d32c0 Documentation for sqlite-utils bulk, refs #377 2022-01-10 17:59:56 -08:00
Simon Willison
71e546e105 Tests for sqlite-utils bulk, refs #377 2022-01-10 17:53:09 -08:00
Simon Willison
de29a10c49 Refactor import_options and insert_upsert_options, refs #377 2022-01-10 17:47:04 -08:00
Simon Willison
d138a60deb --analyze option for create-index, insert, update commands, closes #379, closes #365 2022-01-10 17:40:55 -08:00
Simon Willison
0dbf4b2775 sqlite-utils analyze command, refs #379 2022-01-10 17:40:55 -08:00
Simon Willison
cf73818136 delete_where(analyze=True), closes #378 2022-01-10 17:40:55 -08:00
Simon Willison
131ed27ab0 analyze=True for insert_all/upsert_all, refs #378 2022-01-10 17:40:55 -08:00
Simon Willison
5881a80be1 Improved test_create_index_analyze test, refs #378 2022-01-10 17:40:55 -08:00
Simon Willison
ced9034383 table.create_index(..., analyze=True), refs #378 2022-01-10 17:40:55 -08:00
Simon Willison
04e319a18d db.analyze() and table.analyze() methods, refs #366 2022-01-10 17:40:55 -08:00
Simon Willison
bb16f52681 Initial sqlite-utils bulk prototype, refs #375 2022-01-09 21:31:41 -08:00
12 changed files with 546 additions and 54 deletions

View file

@ -2,6 +2,27 @@
Changelog Changelog
=========== ===========
.. _v3_21:
3.21 (2022-01-10)
-----------------
CLI and Python library improvements to help run `ANALYZE <https://www.sqlite.org/lang_analyze.html>`__ after creating indexes or inserting rows, to gain better performance from the SQLite query planner when it runs against indexes.
Three new CLI commands: ``create-database``, ``analyze`` and ``bulk``.
- New ``sqlite-utils create-database`` command for creating new empty database files. (:issue:`348`)
- New Python methods for running ``ANALYZE`` against a database, table or index: ``db.analyze()`` and ``table.analyze()``, see :ref:`python_api_analyze`. (:issue:`366`)
- New :ref:`sqlite-utils analyze command <cli_analyze>` for running ``ANALYZE`` using the CLI. (:issue:`379`)
- The ``create-index``, ``insert`` and ``update`` commands now have a new ``--analyze`` option for running ``ANALYZE`` after the command has completed. (:issue:`379`)
- New :ref:`sqlite-utils bulk command <cli_bulk>` which can import records in the same way as ``sqlite-utils insert`` (from JSON, CSV or TSV) and use them to bulk execute a parametrized SQL query. (:issue:`375`)
- The CLI tool can now also be run using ``python -m sqlite_utils``. (:issue:`368`)
- Using ``--fmt`` now implies ``--table``, so you don't need to pass both options. (:issue:`374`)
- The ``--convert`` function applied to rows can now modify the row in place. (:issue:`371`)
- The :ref:`insert-files command <cli_insert_files>` supports two new columns: ``stem`` and ``suffix``. (:issue:`372`)
- The ``--nl`` import option now ignores blank lines in the input. (:issue:`376`)
- Fixed bug where streaming input to the ``insert`` command with ``--batch-size 1`` would appear to only commit after several rows had been ingested, due to unnecessary input buffering. (:issue:`364`)
.. _v3_20: .. _v3_20:
3.20 (2022-01-05) 3.20 (2022-01-05)

View file

@ -761,6 +761,8 @@ You can delete all the existing rows in the table before inserting the new recor
$ sqlite-utils insert dogs.db dogs dogs.json --truncate $ sqlite-utils insert dogs.db dogs dogs.json --truncate
You can add the ``--analyze`` option to run ``ANALYZE`` against the table after the rows have been inserted.
.. _cli_inserting_data_binary: .. _cli_inserting_data_binary:
Inserting binary data Inserting binary data
@ -1076,6 +1078,36 @@ The command will fail if you reference columns that do not exist on the table. T
.. note:: .. note::
``upsert`` in sqlite-utils 1.x worked like ``insert ... --replace`` does in 2.x. See `issue #66 <https://github.com/simonw/sqlite-utils/issues/66>`__ for details of this change. ``upsert`` in sqlite-utils 1.x worked like ``insert ... --replace`` does in 2.x. See `issue #66 <https://github.com/simonw/sqlite-utils/issues/66>`__ for details of this change.
.. _cli_bulk:
Executing SQL in bulk
=====================
If you have a JSON, newline-delimited JSON, CSV or TSV file you can execute a bulk SQL query using each of the records in that file using the ``sqlite-utils bulk`` command.
The command takes the database file, the SQL to be executed and the file containing records to be used when evaluating the SQL query.
The SQL query should include ``:named`` parameters that match the keys in the records.
For example, given a ``chickens.csv`` CSV file containing the following::
id,name
1,Blue
2,Snowy
3,Azi
4,Lila
5,Suna
6,Cardi
You could insert those rows into a pre-created ``chickens`` table like so::
$ sqlite-utils bulk chickens.db \
'insert into chickens (id, name) values (:id, :name)' \
chickens.csv --csv
This command takes the same options as the ``sqlite-utils insert`` command - so it defaults to expecting JSON but can accept other formats using ``--csv`` or ``--tsv`` or ``--nl`` or other options described above.
.. _cli_insert_files: .. _cli_insert_files:
Inserting data from files Inserting data from files
@ -1697,6 +1729,8 @@ This will create an index on that table on ``(col1, col2 desc, col3)``.
If your column names are already prefixed with a hyphen you'll need to manually execute a ``CREATE INDEX`` SQL statement to add indexes to them rather than using this tool. If your column names are already prefixed with a hyphen you'll need to manually execute a ``CREATE INDEX`` SQL statement to add indexes to them rather than using this tool.
Add the ``--analyze`` option to run ``ANALYZE`` against the index after it has been created.
.. _cli_fts: .. _cli_fts:
Configuring full-text search Configuring full-text search
@ -1800,6 +1834,25 @@ If the ``_counts`` table ever becomes out-of-sync with the actual table counts y
$ sqlite-utils reset-counts mydb.db $ sqlite-utils reset-counts mydb.db
.. _cli_analyze:
Optimizing index usage with ANALYZE
===================================
The `SQLite ANALYZE command <https://www.sqlite.org/lang_analyze.html>`__ builds a table of statistics which the query planner can use to make better decisions about which indexes to use for a given query.
You should run ``ANALYZE`` if your database is large and you do not think your indexes are being efficiently used.
To run ``ANALYZE`` against every index in a database, use this::
$ sqlite-utils analyze mydb.db
You can run it against specific tables, or against specific named indexes, by passing them as optional arguments::
$ sqlite-utils analyze mydb.db mytable idx_mytable_name
You can also run ``ANALYZE`` as part of another command using the ``--analyze`` option. This is supported by the ``create-index``, ``insert`` and ``upsert`` commands.
.. _cli_vacuum: .. _cli_vacuum:
Vacuum Vacuum

View file

@ -638,8 +638,9 @@ The function can accept an iterator or generator of rows and will commit them ac
You can skip inserting any records that have a primary key that already exists using ``ignore=True``. This works with both ``.insert({...}, ignore=True)`` and ``.insert_all([...], ignore=True)``. You can skip inserting any records that have a primary key that already exists using ``ignore=True``. This works with both ``.insert({...}, ignore=True)`` and ``.insert_all([...], ignore=True)``.
You can delete all the existing rows in the table before inserting the new You can delete all the existing rows in the table before inserting the new records using ``truncate=True``. This is useful if you want to replace the data in the table.
records using ``truncate=True``. This is useful if you want to replace the data in the table.
Pass ``analyze=True`` to run ``ANALYZE`` against the table after inserting the new records.
.. _python_api_insert_replace: .. _python_api_insert_replace:
@ -716,6 +717,8 @@ You can delete all records in a table that match a specific WHERE statement usin
Calling ``table.delete_where()`` with no other arguments will delete every row in the table. Calling ``table.delete_where()`` with no other arguments will delete every row in the table.
Pass ``analyze=True`` to run ``ANALYZE`` against the table after deleting the rows.
.. _python_api_upsert: .. _python_api_upsert:
Upserting data Upserting data
@ -2204,6 +2207,35 @@ You can create a unique index by passing ``unique=True``:
Use ``if_not_exists=True`` to do nothing if an index with that name already exists. Use ``if_not_exists=True`` to do nothing if an index with that name already exists.
Pass ``analyze=True`` to run ``ANALYZE`` against the new index after creating it.
.. _python_api_analyze:
Optimizing index usage with ANALYZE
===================================
The `SQLite ANALYZE command <https://www.sqlite.org/lang_analyze.html>`__ builds a table of statistics which the query planner can use to make better decisions about which indexes to use for a given query.
You should run ``ANALYZE`` if your database is large and you do not think your indexes are being efficiently used.
To run ``ANALYZE`` against every index in a database, use this:
.. code-block:: python
db.analyze()
To run it just against a specific named index, pass the name of the index to that method:
.. code-block:: python
db.analyze("idx_countries_country_name")
To run against all indexes attached to a specific table, you can either pass the table name to ``db.analyze(...)`` or you can call the method directly on the table, like this:
.. code-block:: python
db["dogs"].analyze()
.. _python_api_vacuum: .. _python_api_vacuum:
Vacuum Vacuum

View file

@ -2,7 +2,7 @@ from setuptools import setup, find_packages
import io import io
import os import os
VERSION = "3.20" VERSION = "3.21"
def get_long_description(): def get_long_description():

View file

@ -304,6 +304,26 @@ def rebuild_fts(path, tables, load_extension):
db[table].rebuild_fts() db[table].rebuild_fts()
@cli.command()
@click.argument(
"path",
type=click.Path(exists=True, file_okay=True, dir_okay=False, allow_dash=False),
required=True,
)
@click.argument("names", nargs=-1)
def analyze(path, names):
"""Run ANALYZE against the whole database, or against specific named indexes and tables"""
db = sqlite_utils.Database(path)
try:
if names:
for name in names:
db.analyze(name)
else:
db.analyze()
except sqlite3.OperationalError as e:
raise click.ClickException(e)
@cli.command() @cli.command()
@click.argument( @click.argument(
"path", "path",
@ -470,8 +490,15 @@ def index_foreign_keys(path, load_extension):
default=False, default=False,
is_flag=True, is_flag=True,
) )
@click.option(
"--analyze",
help="Run ANALYZE after creating the index",
is_flag=True,
)
@load_extension_option @load_extension_option
def create_index(path, table, column, name, unique, if_not_exists, load_extension): def create_index(
path, table, column, name, unique, if_not_exists, analyze, load_extension
):
""" """
Add an index to the specified table covering the specified columns. Add an index to the specified table covering the specified columns.
Use "sqlite-utils create-index mydb -- -column" to specify descending Use "sqlite-utils create-index mydb -- -column" to specify descending
@ -486,7 +513,11 @@ def create_index(path, table, column, name, unique, if_not_exists, load_extensio
col = DescIndex(col[1:]) col = DescIndex(col[1:])
columns.append(col) columns.append(col)
db[table].create_index( db[table].create_index(
columns, index_name=name, unique=unique, if_not_exists=if_not_exists columns,
index_name=name,
unique=unique,
if_not_exists=if_not_exists,
analyze=analyze,
) )
@ -629,19 +660,7 @@ def reset_counts(path, load_extension):
db.reset_counts() db.reset_counts()
def insert_upsert_options(fn): _import_options = (
for decorator in reversed(
(
click.argument(
"path",
type=click.Path(file_okay=True, dir_okay=False, allow_dash=False),
required=True,
),
click.argument("table"),
click.argument("file", type=click.File("rb"), required=True),
click.option(
"--pk", help="Columns to use as the primary key, e.g. id", multiple=True
),
click.option( click.option(
"--flatten", "--flatten",
is_flag=True, is_flag=True,
@ -670,12 +689,37 @@ def insert_upsert_options(fn):
), ),
click.option("--delimiter", help="Delimiter to use for CSV files"), click.option("--delimiter", help="Delimiter to use for CSV files"),
click.option("--quotechar", help="Quote character to use for CSV/TSV"), click.option("--quotechar", help="Quote character to use for CSV/TSV"),
click.option("--sniff", is_flag=True, help="Detect delimiter and quote character"),
click.option("--no-headers", is_flag=True, help="CSV file has no header row"),
click.option( click.option(
"--sniff", is_flag=True, help="Detect delimiter and quote character" "--encoding",
help="Character encoding for input, defaults to utf-8",
), ),
)
def import_options(fn):
for decorator in reversed(_import_options):
fn = decorator(fn)
return fn
def insert_upsert_options(fn):
for decorator in reversed(
(
click.argument(
"path",
type=click.Path(file_okay=True, dir_okay=False, allow_dash=False),
required=True,
),
click.argument("table"),
click.argument("file", type=click.File("rb"), required=True),
click.option( click.option(
"--no-headers", is_flag=True, help="CSV file has no header row" "--pk", help="Columns to use as the primary key, e.g. id", multiple=True
), ),
)
+ _import_options
+ (
click.option( click.option(
"--batch-size", type=int, default=100, help="Commit every X records" "--batch-size", type=int, default=100, help="Commit every X records"
), ),
@ -695,10 +739,6 @@ def insert_upsert_options(fn):
type=(str, str), type=(str, str),
help="Default value that should be set for a column", help="Default value that should be set for a column",
), ),
click.option(
"--encoding",
help="Character encoding for input, defaults to utf-8",
),
click.option( click.option(
"-d", "-d",
"--detect-types", "--detect-types",
@ -706,6 +746,11 @@ def insert_upsert_options(fn):
envvar="SQLITE_UTILS_DETECT_TYPES", envvar="SQLITE_UTILS_DETECT_TYPES",
help="Detect types for columns in CSV/TSV data", help="Detect types for columns in CSV/TSV data",
), ),
click.option(
"--analyze",
is_flag=True,
help="Run ANALYZE at the end of this operation",
),
load_extension_option, load_extension_option,
click.option("--silent", is_flag=True, help="Do not show progress bar"), click.option("--silent", is_flag=True, help="Do not show progress bar"),
) )
@ -731,6 +776,7 @@ def insert_upsert_implementation(
quotechar, quotechar,
sniff, sniff,
no_headers, no_headers,
encoding,
batch_size, batch_size,
alter, alter,
upsert, upsert,
@ -739,10 +785,11 @@ def insert_upsert_implementation(
truncate=False, truncate=False,
not_null=None, not_null=None,
default=None, default=None,
encoding=None,
detect_types=None, detect_types=None,
analyze=False,
load_extension=None, load_extension=None,
silent=False, silent=False,
bulk_sql=None,
): ):
db = sqlite_utils.Database(path) db = sqlite_utils.Database(path)
_load_extensions(db, load_extension) _load_extensions(db, load_extension)
@ -833,7 +880,12 @@ def insert_upsert_implementation(
else: else:
docs = (fn(doc) or doc for doc in docs) docs = (fn(doc) or doc for doc in docs)
extra_kwargs = {"ignore": ignore, "replace": replace, "truncate": truncate} extra_kwargs = {
"ignore": ignore,
"replace": replace,
"truncate": truncate,
"analyze": analyze,
}
if not_null: if not_null:
extra_kwargs["not_null"] = set(not_null) extra_kwargs["not_null"] = set(not_null)
if default: if default:
@ -844,6 +896,12 @@ def insert_upsert_implementation(
# Apply {"$base64": true, ...} decoding, if needed # Apply {"$base64": true, ...} decoding, if needed
docs = (decode_base64_values(doc) for doc in docs) docs = (decode_base64_values(doc) for doc in docs)
# For bulk_sql= we use cursor.executemany() instead
if bulk_sql:
with db.conn:
db.conn.cursor().executemany(bulk_sql, docs)
return
try: try:
db[table].insert_all( db[table].insert_all(
docs, pk=pk, batch_size=batch_size, alter=alter, **extra_kwargs docs, pk=pk, batch_size=batch_size, alter=alter, **extra_kwargs
@ -926,10 +984,11 @@ def insert(
quotechar, quotechar,
sniff, sniff,
no_headers, no_headers,
encoding,
batch_size, batch_size,
alter, alter,
encoding,
detect_types, detect_types,
analyze,
load_extension, load_extension,
silent, silent,
ignore, ignore,
@ -977,14 +1036,15 @@ def insert(
quotechar, quotechar,
sniff, sniff,
no_headers, no_headers,
encoding,
batch_size, batch_size,
alter=alter, alter=alter,
upsert=False, upsert=False,
ignore=ignore, ignore=ignore,
replace=replace, replace=replace,
truncate=truncate, truncate=truncate,
encoding=encoding,
detect_types=detect_types, detect_types=detect_types,
analyze=analyze,
load_extension=load_extension, load_extension=load_extension,
silent=silent, silent=silent,
not_null=not_null, not_null=not_null,
@ -1014,11 +1074,12 @@ def upsert(
quotechar, quotechar,
sniff, sniff,
no_headers, no_headers,
encoding,
alter, alter,
not_null, not_null,
default, default,
encoding,
detect_types, detect_types,
analyze,
load_extension, load_extension,
silent, silent,
): ):
@ -1045,13 +1106,14 @@ def upsert(
quotechar, quotechar,
sniff, sniff,
no_headers, no_headers,
encoding,
batch_size, batch_size,
alter=alter, alter=alter,
upsert=True, upsert=True,
not_null=not_null, not_null=not_null,
default=default, default=default,
encoding=encoding,
detect_types=detect_types, detect_types=detect_types,
analyze=analyze,
load_extension=load_extension, load_extension=load_extension,
silent=silent, silent=silent,
) )
@ -1059,6 +1121,71 @@ def upsert(
raise click.ClickException(UNICODE_ERROR.format(ex)) raise click.ClickException(UNICODE_ERROR.format(ex))
@cli.command()
@click.argument(
"path",
type=click.Path(file_okay=True, dir_okay=False, allow_dash=False),
required=True,
)
@click.argument("sql")
@click.argument("file", type=click.File("rb"), required=True)
@import_options
@load_extension_option
def bulk(
path,
file,
sql,
flatten,
nl,
csv,
tsv,
lines,
text,
convert,
imports,
delimiter,
quotechar,
sniff,
no_headers,
encoding,
load_extension,
):
"""
Execute parameterized SQL against the provided list of documents.
"""
try:
insert_upsert_implementation(
path=path,
table=None,
file=file,
pk=None,
flatten=flatten,
nl=nl,
csv=csv,
tsv=tsv,
lines=lines,
text=text,
convert=convert,
imports=imports,
delimiter=delimiter,
quotechar=quotechar,
sniff=sniff,
no_headers=no_headers,
encoding=encoding,
batch_size=1,
alter=False,
upsert=False,
not_null=set(),
default={},
detect_types=False,
load_extension=load_extension,
silent=False,
bulk_sql=sql,
)
except (sqlite3.OperationalError, sqlite3.IntegrityError) as e:
raise click.ClickException(str(e))
@cli.command(name="create-database") @cli.command(name="create-database")
@click.argument( @click.argument(
"path", "path",

View file

@ -923,6 +923,13 @@ class Database:
"Run a SQLite ``VACUUM`` against the database." "Run a SQLite ``VACUUM`` against the database."
self.execute("VACUUM;") self.execute("VACUUM;")
def analyze(self, name=None):
"Run ``ANALYZE`` against the entire database or a named table or index."
sql = "ANALYZE"
if name is not None:
sql += " [{}]".format(name)
self.execute(sql)
class Queryable: class Queryable:
def exists(self) -> bool: def exists(self) -> bool:
@ -1547,6 +1554,7 @@ class Table(Queryable):
unique: bool = False, unique: bool = False,
if_not_exists: bool = False, if_not_exists: bool = False,
find_unique_name: bool = False, find_unique_name: bool = False,
analyze: bool = False,
): ):
""" """
Create an index on this table. Create an index on this table.
@ -1558,6 +1566,7 @@ class Table(Queryable):
- ``if_not_exists`` - only create the index if one with that name does not already exist. - ``if_not_exists`` - only create the index if one with that name does not already exist.
- ``find_unique_name`` - if ``index_name`` is not provided and the automatically derived name - ``find_unique_name`` - if ``index_name`` is not provided and the automatically derived name
already exists, keep incrementing a suffix number to find an available name. already exists, keep incrementing a suffix number to find an available name.
- ``analyze`` - run ``ANALYZE`` against this index after creating it.
See :ref:`python_api_create_index`. See :ref:`python_api_create_index`.
""" """
@ -1574,7 +1583,11 @@ class Table(Queryable):
columns_sql.append(fmt.format(column)) columns_sql.append(fmt.format(column))
suffix = None suffix = None
created_index_name = None
while True: while True:
created_index_name = (
"{}_{}".format(index_name, suffix) if suffix else index_name
)
sql = ( sql = (
textwrap.dedent( textwrap.dedent(
""" """
@ -1584,9 +1597,7 @@ class Table(Queryable):
) )
.strip() .strip()
.format( .format(
index_name="{}_{}".format(index_name, suffix) index_name=created_index_name,
if suffix
else index_name,
table_name=self.name, table_name=self.name,
columns=", ".join(columns_sql), columns=", ".join(columns_sql),
unique="UNIQUE " if unique else "", unique="UNIQUE " if unique else "",
@ -1611,6 +1622,8 @@ class Table(Queryable):
continue continue
else: else:
raise e raise e
if analyze:
self.db.analyze(created_index_name)
return self return self
def add_column( def add_column(
@ -2106,15 +2119,28 @@ class Table(Queryable):
return self return self
def delete_where( def delete_where(
self, where: str = None, where_args: Optional[Union[Iterable, dict]] = None self,
where: str = None,
where_args: Optional[Union[Iterable, dict]] = None,
analyze: bool = False,
) -> "Table": ) -> "Table":
"Delete rows matching specified where clause, or delete all rows in the table." """
Delete rows matching the specified where clause, or delete all rows in the table.
- ``where`` - a SQL fragment to use as a ``WHERE`` clause, for example ``age > ?`` or ``age > :age``.
- ``where_args`` - a list of arguments (if using ``?``) or a dictionary (if using ``:age``).
- ``analyze`` - set to ``True`` to run ``ANALYZE`` after the rows have been deleted.
See :ref:`python_api_delete_where`.
"""
if not self.exists(): if not self.exists():
return self return self
sql = "delete from [{}]".format(self.name) sql = "delete from [{}]".format(self.name)
if where is not None: if where is not None:
sql += " where " + where sql += " where " + where
self.db.execute(sql, where_args or []) self.db.execute(sql, where_args or [])
if analyze:
self.analyze()
return self return self
def update( def update(
@ -2562,10 +2588,13 @@ class Table(Queryable):
conversions=DEFAULT, conversions=DEFAULT,
columns=DEFAULT, columns=DEFAULT,
upsert=False, upsert=False,
analyze=False,
) -> "Table": ) -> "Table":
""" """
Like ``.insert()`` but takes a list of records and ensures that the table Like ``.insert()`` but takes a list of records and ensures that the table
that it creates (if table does not exist) has columns for ALL of that data. that it creates (if table does not exist) has columns for ALL of that data.
Use ``analyze=True`` to run ``ANALYZE`` after the insert has completed.
""" """
pk = self.value_or_default("pk", pk) pk = self.value_or_default("pk", pk)
foreign_keys = self.value_or_default("foreign_keys", foreign_keys) foreign_keys = self.value_or_default("foreign_keys", foreign_keys)
@ -2658,6 +2687,9 @@ class Table(Queryable):
ignore, ignore,
) )
if analyze:
self.analyze()
return self return self
def upsert( def upsert(
@ -2708,6 +2740,7 @@ class Table(Queryable):
extracts=DEFAULT, extracts=DEFAULT,
conversions=DEFAULT, conversions=DEFAULT,
columns=DEFAULT, columns=DEFAULT,
analyze=False,
) -> "Table": ) -> "Table":
""" """
Like ``.upsert()`` but can be applied to a list of records. Like ``.upsert()`` but can be applied to a list of records.
@ -2726,6 +2759,7 @@ class Table(Queryable):
conversions=conversions, conversions=conversions,
columns=columns, columns=columns,
upsert=True, upsert=True,
analyze=analyze,
) )
def add_missing_columns(self, records: Iterable[Dict[str, Any]]) -> "Table": def add_missing_columns(self, records: Iterable[Dict[str, Any]]) -> "Table":
@ -2902,6 +2936,10 @@ class Table(Queryable):
) )
return self return self
def analyze(self):
"Run ANALYZE against this table"
self.db.analyze(self.name)
def analyze_column( def analyze_column(
self, column: str, common_limit: int = 10, value_truncate=None, total_rows=None self, column: str, common_limit: int = 10, value_truncate=None, total_rows=None
) -> "ColumnDetails": ) -> "ColumnDetails":

45
tests/test_analyze.py Normal file
View file

@ -0,0 +1,45 @@
import pytest
@pytest.fixture
def db(fresh_db):
fresh_db["one_index"].insert({"id": 1, "name": "Cleo"}, pk="id")
fresh_db["one_index"].create_index(["name"])
fresh_db["two_indexes"].insert({"id": 1, "name": "Cleo", "species": "dog"}, pk="id")
fresh_db["two_indexes"].create_index(["name"])
fresh_db["two_indexes"].create_index(["species"])
return fresh_db
def test_analyze_whole_database(db):
assert set(db.table_names()) == {"one_index", "two_indexes"}
db.analyze()
assert set(db.table_names()) == {"one_index", "two_indexes", "sqlite_stat1"}
assert list(db["sqlite_stat1"].rows) == [
{"tbl": "two_indexes", "idx": "idx_two_indexes_species", "stat": "1 1"},
{"tbl": "two_indexes", "idx": "idx_two_indexes_name", "stat": "1 1"},
{"tbl": "one_index", "idx": "idx_one_index_name", "stat": "1 1"},
]
@pytest.mark.parametrize("method", ("db_method_with_name", "table_method"))
def test_analyze_one_table(db, method):
assert set(db.table_names()) == {"one_index", "two_indexes"}
if method == "db_method_with_name":
db.analyze("one_index")
elif method == "table_method":
db["one_index"].analyze()
assert set(db.table_names()) == {"one_index", "two_indexes", "sqlite_stat1"}
assert list(db["sqlite_stat1"].rows) == [
{"tbl": "one_index", "idx": "idx_one_index_name", "stat": "1 1"}
]
def test_analyze_index_by_name(db):
assert set(db.table_names()) == {"one_index", "two_indexes"}
db.analyze("idx_two_indexes_species")
assert set(db.table_names()) == {"one_index", "two_indexes", "sqlite_stat1"}
assert list(db["sqlite_stat1"].rows) == [
{"tbl": "two_indexes", "idx": "idx_two_indexes_species", "stat": "1 1"},
]

View file

@ -224,6 +224,17 @@ def test_create_index(db_path):
) )
def test_create_index_analyze(db_path):
db = Database(db_path)
assert "sqlite_stat1" not in db.table_names()
assert [] == db["Gosh"].indexes
result = CliRunner().invoke(
cli.cli, ["create-index", db_path, "Gosh", "c1", "--analyze"]
)
assert result.exit_code == 0
assert "sqlite_stat1" in db.table_names()
def test_create_index_desc(db_path): def test_create_index_desc(db_path):
db = Database(db_path) db = Database(db_path)
assert [] == db["Gosh"].indexes assert [] == db["Gosh"].indexes
@ -889,6 +900,20 @@ def test_upsert(db_path, tmpdir):
] ]
def test_upsert_analyze(db_path, tmpdir):
db = Database(db_path)
db["rows"].insert({"id": 1, "foo": "x", "n": 3}, pk="id")
db["rows"].create_index(["n"])
assert "sqlite_stat1" not in db.table_names()
result = CliRunner().invoke(
cli.cli,
["upsert", db_path, "rows", "-", "--nl", "--analyze", "--pk", "id"],
input='{"id": 2, "foo": "bar", "n": 1}',
)
assert 0 == result.exit_code, result.output
assert "sqlite_stat1" in db.table_names()
def test_upsert_flatten(tmpdir): def test_upsert_flatten(tmpdir):
db_path = str(tmpdir / "flat.db") db_path = str(tmpdir / "flat.db")
db = Database(db_path) db = Database(db_path)
@ -2057,3 +2082,41 @@ def test_create_database(tmpdir, enable_wal):
assert db.journal_mode == "wal" assert db.journal_mode == "wal"
else: else:
assert db.journal_mode == "delete" assert db.journal_mode == "delete"
@pytest.mark.parametrize(
"options,expected",
(
(
[],
[
{"tbl": "two_indexes", "idx": "idx_two_indexes_species", "stat": "1 1"},
{"tbl": "two_indexes", "idx": "idx_two_indexes_name", "stat": "1 1"},
{"tbl": "one_index", "idx": "idx_one_index_name", "stat": "1 1"},
],
),
(
["one_index"],
[
{"tbl": "one_index", "idx": "idx_one_index_name", "stat": "1 1"},
],
),
(
["idx_two_indexes_name"],
[
{"tbl": "two_indexes", "idx": "idx_two_indexes_name", "stat": "1 1"},
],
),
),
)
def test_analyze(tmpdir, options, expected):
db_path = str(tmpdir / "test.db")
db = Database(db_path)
db["one_index"].insert({"id": 1, "name": "Cleo"}, pk="id")
db["one_index"].create_index(["name"])
db["two_indexes"].insert({"id": 1, "name": "Cleo", "species": "dog"}, pk="id")
db["two_indexes"].create_index(["name"])
db["two_indexes"].create_index(["species"])
result = CliRunner().invoke(cli.cli, ["analyze", db_path] + options)
assert result.exit_code == 0
assert list(db["sqlite_stat1"].rows) == expected

57
tests/test_cli_bulk.py Normal file
View file

@ -0,0 +1,57 @@
from click.testing import CliRunner
from sqlite_utils import cli, Database
import pathlib
import pytest
@pytest.fixture
def test_db_and_path(tmpdir):
db_path = str(pathlib.Path(tmpdir) / "data.db")
db = Database(db_path)
db["example"].insert_all(
[
{"id": 1, "name": "One"},
{"id": 2, "name": "Two"},
],
pk="id",
)
return db, db_path
def test_cli_bulk(test_db_and_path):
db, db_path = test_db_and_path
result = CliRunner().invoke(
cli.cli,
[
"bulk",
db_path,
"insert into example (id, name) values (:id, :name)",
"-",
"--nl",
],
input='{"id": 3, "name": "Three"}\n{"id": 4, "name": "Four"}\n',
)
assert result.exit_code == 0, result.output
assert [
{"id": 1, "name": "One"},
{"id": 2, "name": "Two"},
{"id": 3, "name": "Three"},
{"id": 4, "name": "Four"},
] == list(db["example"].rows)
def test_cli_bulk_error(test_db_and_path):
_, db_path = test_db_and_path
result = CliRunner().invoke(
cli.cli,
[
"bulk",
db_path,
"insert into example (id, name) value (:id, :name)",
"-",
"--nl",
],
input='{"id": 3, "name": "Three"}',
)
assert result.exit_code == 1
assert result.output == 'Error: near "value": syntax error\n'

View file

@ -327,6 +327,20 @@ def test_insert_alter(db_path, tmpdir):
] == list(db.query("select foo, n, baz from from_json_nl")) ] == list(db.query("select foo, n, baz from from_json_nl"))
def test_insert_analyze(db_path):
db = Database(db_path)
db["rows"].insert({"foo": "x", "n": 3})
db["rows"].create_index(["n"])
assert "sqlite_stat1" not in db.table_names()
result = CliRunner().invoke(
cli.cli,
["insert", db_path, "rows", "-", "--nl", "--analyze"],
input='{"foo": "bar", "n": 1}\n{"foo": "baz", "n": 2}',
)
assert 0 == result.exit_code, result.output
assert "sqlite_stat1" in db.table_names()
def test_insert_lines(db_path): def test_insert_lines(db_path):
result = CliRunner().invoke( result = CliRunner().invoke(
cli.cli, cli.cli,

View file

@ -774,6 +774,17 @@ def test_create_index_find_unique_name(fresh_db):
assert index_names == {"idx_t_id", "idx_t_id_2", "idx_t_id_3"} assert index_names == {"idx_t_id", "idx_t_id_2", "idx_t_id_3"}
def test_create_index_analyze(fresh_db):
dogs = fresh_db["dogs"]
assert "sqlite_stat1" not in fresh_db.table_names()
dogs.insert({"name": "Cleo", "twitter": "cleopaws"})
dogs.create_index(["name"], analyze=True)
assert "sqlite_stat1" in fresh_db.table_names()
assert list(fresh_db["sqlite_stat1"].rows) == [
{"tbl": "dogs", "idx": "idx_dogs_name", "stat": "1 1"}
]
@pytest.mark.parametrize( @pytest.mark.parametrize(
"data_structure", "data_structure",
( (
@ -1017,6 +1028,23 @@ def test_insert_all_single_column(fresh_db):
assert table.pks == ["name"] assert table.pks == ["name"]
@pytest.mark.parametrize("method_name", ("insert_all", "upsert_all"))
def test_insert_all_analyze(fresh_db, method_name):
table = fresh_db["table"]
table.insert_all([{"id": 1, "name": "Cleo"}], pk="id")
assert "sqlite_stat1" not in fresh_db.table_names()
table.create_index(["name"], analyze=True)
assert list(fresh_db["sqlite_stat1"].rows) == [
{"tbl": "table", "idx": "idx_table_name", "stat": "1 1"}
]
method = getattr(table, method_name)
method([{"id": 2, "name": "Suna"}], pk="id", analyze=True)
assert "sqlite_stat1" in fresh_db.table_names()
assert list(fresh_db["sqlite_stat1"].rows) == [
{"tbl": "table", "idx": "idx_table_name", "stat": "2 1"}
]
def test_create_with_a_null_column(fresh_db): def test_create_with_a_null_column(fresh_db):
record = {"name": "Name", "description": None} record = {"name": "Name", "description": None}
fresh_db["t"].insert(record) fresh_db["t"].insert(record)

View file

@ -30,3 +30,17 @@ def test_delete_where_all(fresh_db):
assert 10 == table.count assert 10 == table.count
table.delete_where() table.delete_where()
assert 0 == table.count assert 0 == table.count
def test_delete_where_analyze(fresh_db):
table = fresh_db["table"]
table.insert_all(({"id": i, "i": i} for i in range(10)), pk="id")
table.create_index(["i"], analyze=True)
assert "sqlite_stat1" in fresh_db.table_names()
assert list(fresh_db["sqlite_stat1"].rows) == [
{"tbl": "table", "idx": "idx_table_i", "stat": "10 1"}
]
table.delete_where("id > ?", [5], analyze=True)
assert list(fresh_db["sqlite_stat1"].rows) == [
{"tbl": "table", "idx": "idx_table_i", "stat": "6 1"}
]