mirror of
https://github.com/simonw/datasette.git
synced 2026-09-28 04:44:21 +02:00
Three corrections to the emitted data, bundled because changing what is on the wire after operators have built dashboards on it is a breaking change - so they belong in the first release that ships spans at all, not a later one. db.query is now SpanKind.CLIENT. Trace UIs key their database rendering off the span kind rather than off db.system, so the spans rendered as ordinary internal work despite carrying db.system and db.query.text. The three child spans stay INTERNAL on purpose: db.query.execute, db.write.execute and db.write.queue_wait are Datasette's decomposition of one logical query, not three database calls, and queue_wait touches no database at all - marking them CLIENT would make one query look like several to anything counting spans by kind. The instrumentation scope now carries the Datasette version and a schema URL, so a backend can tell which Datasette produced a span. The URL is 1.29.0 rather than the latest semconv release because that is the highest version at which every name emitted here is the current spelling: db.system was renamed to db.system.name in 1.30.0 and this code still emits the older form. Claiming a later schema would be false, and would stop a consumer translating that name forward, since the claim asserts the rename already happened. db.operation.name is the statement's leading keyword matched against a fixed allowlist, not a parse. On a public instance the SQL is attacker-controlled and this attribute is a candidate metric dimension in a later phase, so echoing back an arbitrary first token would let a visitor's typo mint a permanent series. Anything unrecognised gets no attribute rather than a wrong one. execute_write_script() does not set it at all, since semantic conventions say not to extract an operation name from query text that can hold several statements. db.collection.name comes only from a new table= argument on Database.execute(), and is never derived from the SQL: deriving it would be a parse, and on an instance where anyone can create a table the value set has no ceiling. It is passed from every query in the table and row views that targets exactly one user table. Internal-catalog reads and the row view's cross-table foreign key counts are deliberately left without it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
111 lines
4.3 KiB
Python
111 lines
4.3 KiB
Python
"""
|
|
OpenTelemetry integration for Datasette core.
|
|
|
|
Core depends on `opentelemetry-api` only. It never creates a
|
|
`TracerProvider`, never configures an exporter, and never touches
|
|
sampling - that is the responsibility of whoever is running Datasette
|
|
(an `opentelemetry-instrument` agent, a future plugin, or a test
|
|
harness). With no provider installed every span produced here is a
|
|
`NonRecordingSpan` and costs approximately nothing.
|
|
"""
|
|
|
|
import re
|
|
|
|
from opentelemetry import trace as otel_trace
|
|
|
|
from .version import __version__
|
|
|
|
# The semantic-convention version whose spellings this instrumentation
|
|
# actually emits. Deliberately NOT the latest release.
|
|
#
|
|
# A schema URL is a machine-readable claim: a consumer doing schema
|
|
# translation replays the renames between the declared version and the one
|
|
# it wants, so the claim has to name the version whose spellings are on the
|
|
# wire. A wrong one makes translation wrong rather than merely uninformative.
|
|
#
|
|
# Datasette emits `db.system`, which was renamed to `db.system.name` in
|
|
# semconv 1.30.0. Everything else it emits (`db.namespace`, `db.query.text`,
|
|
# `db.operation.name`, `db.collection.name`) has been current since 1.26.0.
|
|
# So 1.29.0 is the highest version at which every name emitted here is the
|
|
# current spelling. Everything under `datasette.*` is Datasette's own and
|
|
# outside semconv, so it is unaffected either way.
|
|
#
|
|
# Declaring 1.43.0 would be false about `db.system`, and would actively STOP
|
|
# a consumer translating it forward, because it asserts the rename already
|
|
# happened. Bump this deliberately, in the same commit as the attribute
|
|
# renames it implies - it is a claim about the names, not decoration.
|
|
SCHEMA_URL = "https://opentelemetry.io/schemas/1.29.0"
|
|
|
|
tracer = otel_trace.get_tracer("datasette", __version__, schema_url=SCHEMA_URL)
|
|
|
|
MAX_SQL_LENGTH = 2048
|
|
|
|
|
|
def sql_attribute(sql: str) -> str:
|
|
"Truncate SQL text so it is safe to attach to a span as an attribute."
|
|
sql = sql.strip()
|
|
if len(sql) <= MAX_SQL_LENGTH:
|
|
return sql
|
|
return sql[:MAX_SQL_LENGTH] + "…[truncated]"
|
|
|
|
|
|
# db.operation.name is the leading keyword of a statement matched against a
|
|
# fixed allowlist - deliberately not a parse.
|
|
#
|
|
# This runs against arbitrary user-supplied SQL (the `?sql=` query string,
|
|
# canned queries, anything typed into the query editor), and the attribute is
|
|
# a candidate dimension on a query-duration metric in a later phase. A metric
|
|
# series is keyed by its attribute values, so echoing back an arbitrary first
|
|
# token would let one visitor's typo mint a new, permanent series. The
|
|
# allowlist bounds that at a fixed, small set regardless of what anyone sends.
|
|
DB_OPERATION_ALLOWLIST = frozenset(
|
|
{
|
|
"SELECT",
|
|
"INSERT",
|
|
"UPDATE",
|
|
"DELETE",
|
|
"CREATE",
|
|
"DROP",
|
|
"ALTER",
|
|
"PRAGMA",
|
|
"EXPLAIN",
|
|
"REPLACE",
|
|
"VACUUM",
|
|
"ANALYZE",
|
|
"WITH",
|
|
}
|
|
)
|
|
|
|
_LEADING_KEYWORD = re.compile(r"^\s*([A-Za-z]+)")
|
|
|
|
|
|
def sql_operation_name(sql: str) -> str | None:
|
|
"""
|
|
The statement's leading keyword, if it is one we recognise.
|
|
|
|
Returns None - never a guess - for anything not on the allowlist,
|
|
including a statement that opens with a comment or with punctuation such
|
|
as the "(" of a parenthesised SELECT.
|
|
|
|
Known limitation: a statement beginning with a CTE reports `WITH` rather
|
|
than the operation inside it, and a substantial share of Datasette's own
|
|
reads take that form. Extracting more than the leading keyword means
|
|
handling comment stripping, parenthesised `(SELECT ...) UNION` and
|
|
compound names like `CREATE TABLE` - each a special case a hand-rolled
|
|
matcher would accrete and eventually get wrong. Omitting a name beats
|
|
guessing at one.
|
|
|
|
Only safe to call with a single statement: `execute_write_script()` runs
|
|
several separated by semicolons, and semantic conventions say
|
|
`db.operation.name` "SHOULD NOT be extracted from db.query.text, when the
|
|
database system supports query text with multiple operations in non-batch
|
|
operations" - so that call site does not use this at all rather than
|
|
reporting only the first statement's operation.
|
|
"""
|
|
match = _LEADING_KEYWORD.match(sql)
|
|
if not match:
|
|
return None
|
|
keyword = match.group(1).upper()
|
|
if keyword in DB_OPERATION_ALLOWLIST:
|
|
return keyword
|
|
return None
|