mirror of
https://github.com/simonw/datasette.git
synced 2026-09-27 12:24:07 +02:00
Add OpenTelemetry metrics for SQL thread pool saturation and query latency
Spans describe requests that have finished. They structurally cannot answer "am I saturating my 3 SQL threads right now", because that is a level rather than an event - and with num_sql_threads defaulting to 3, it is usually the first thing worth knowing about a busy Datasette. This adds the metrics that answer it. Five observable gauges, computed only when something is collecting, so an instance with no MeterProvider installed does no work for them at all: datasette.sql.threads.limit num_sql_threads datasette.sql.threads.queue_depth queries waiting for a free thread datasette.sql.queries.pending in-flight reads, by db.namespace datasette.write.queue_depth writes behind the single write thread datasette.connections.open tracked file connections Three instruments recorded inline, which matters because metrics survive trace sampling and spans do not - an operator sampling 1% of traces still gets 100% of the latency distribution: db.client.operation.duration semconv histogram, with error.type datasette.write.queue_wait the metric twin of the existing span datasette.sql.queries.interrupted sql_time_limit_ms kills The interrupted counter closes a gap the plan called out as unanswerable: "how often are we killing queries at the limit" is a rate, and a rate cannot be recovered from sampled spans. Core still creates no provider of any kind, so the architecture is unchanged; `grep -rn 'opentelemetry.sdk' datasette/` stays empty. One real difference from tracing is worth recording: _ProxyMeter and its instruments forward to a provider installed after they were created, whereas ProxyTracer permanently caches the first concrete tracer it resolves. Module-level instruments are therefore safe and the test fixture has no ordering constraint. Live instances are tracked in a lock-guarded WeakSet so instrumenting an instance never keeps it alive. The pool gauges carry no attribute saying which Datasette produced them: production runs one instance per process, and adding an id to disambiguate the test suite's hundreds of instances would buy unbounded attribute cardinality to fix a case that does not occur. The collision is documented instead, and the gauge callbacks are plain generator functions so tests can assert exact values by calling them directly rather than through the SDK's last-value aggregation. demos/otel/metrics_demo.py fires 12 concurrent 40ms queries at a 3-thread pool and samples the gauges mid-flight: queue_depth peaks at exactly 9, and the duration histogram reads max=0.1695s for a query whose work is 40ms. That gap is the queue, and it is the thing traces alone will not show you. Also corrects the demo README's privacy section, which still claimed parameter values are never recorded - that stopped being unconditionally true when trace_sql_parameters landed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> (cherry picked from 6ef0dd8c and adapted to the rebuilt phase-1 stack: attribute names now come from telemetry_registry where entries exist, the meter carries the instrumentation-scope version and schema URL, and the interrupted-queries counter skips expected timeouts - callers that opted into a deliberately short budget, like facet suggestion - matching how those are excluded from span error status. The internals.rst reference lands with the registry commit that follows.) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01F2h9ANGZ7paWSpqs5DUAcG
This commit is contained in:
parent
9bce2eec7c
commit
90b727db73
6 changed files with 730 additions and 16 deletions
|
|
@ -114,6 +114,106 @@ def otel_spans():
|
|||
yield _otel_span_exporter
|
||||
|
||||
|
||||
_otel_metric_reader = None
|
||||
|
||||
|
||||
@pytest.fixture(scope="session", autouse=True)
|
||||
def _otel_meter_provider():
|
||||
"""
|
||||
Install a real OTel SDK MeterProvider + InMemoryMetricReader once per
|
||||
process.
|
||||
|
||||
Unlike the tracer, ordering is not load-bearing here: `_ProxyMeter` and
|
||||
the `_ProxyInstrument`s it hands out forward to a provider installed
|
||||
*after* they were created, whereas `ProxyTracer` permanently caches the
|
||||
first concrete tracer it resolves. This fixture is still session-scoped
|
||||
and autouse for symmetry, and so that a single reader collects for the
|
||||
whole run.
|
||||
|
||||
DELTA temporality is chosen for counters and histograms so that each
|
||||
collection reports only what happened since the previous one. With the
|
||||
SDK default of CUMULATIVE, every metrics test would see every query run
|
||||
by every earlier test in the session.
|
||||
"""
|
||||
global _otel_metric_reader
|
||||
try:
|
||||
from opentelemetry import metrics as otel_metrics
|
||||
from opentelemetry.sdk.metrics import Counter, Histogram, MeterProvider
|
||||
from opentelemetry.sdk.metrics.export import (
|
||||
AggregationTemporality,
|
||||
InMemoryMetricReader,
|
||||
)
|
||||
except ImportError:
|
||||
return
|
||||
reader = InMemoryMetricReader(
|
||||
preferred_temporality={
|
||||
Counter: AggregationTemporality.DELTA,
|
||||
Histogram: AggregationTemporality.DELTA,
|
||||
}
|
||||
)
|
||||
otel_metrics.set_meter_provider(MeterProvider(metric_readers=[reader]))
|
||||
_otel_metric_reader = reader
|
||||
|
||||
|
||||
class MetricsCollector:
|
||||
"""
|
||||
Thin reader over an `InMemoryMetricReader`.
|
||||
|
||||
`collect()` runs a collection cycle - which is what invokes the observable
|
||||
gauge callbacks - and snapshots the result. Queries then run against that
|
||||
snapshot rather than re-collecting, so a test that inspects several
|
||||
metrics sees one consistent moment and does not drain delta state twice.
|
||||
"""
|
||||
|
||||
def __init__(self, reader):
|
||||
self.reader = reader
|
||||
self.snapshot = {}
|
||||
|
||||
def collect(self):
|
||||
self.snapshot = {}
|
||||
data = self.reader.get_metrics_data()
|
||||
if data is None:
|
||||
return self.snapshot
|
||||
for resource_metrics in data.resource_metrics:
|
||||
for scope_metrics in resource_metrics.scope_metrics:
|
||||
for metric in scope_metrics.metrics:
|
||||
self.snapshot.setdefault(metric.name, []).extend(
|
||||
metric.data.data_points
|
||||
)
|
||||
return self.snapshot
|
||||
|
||||
def points(self, name, attributes=None):
|
||||
"Data points for `name` whose attributes are a superset of `attributes`."
|
||||
found = []
|
||||
for point in self.snapshot.get(name, []):
|
||||
point_attributes = dict(point.attributes or {})
|
||||
if all(point_attributes.get(k) == v for k, v in (attributes or {}).items()):
|
||||
found.append(point)
|
||||
return found
|
||||
|
||||
def point(self, name, attributes=None):
|
||||
"The single matching data point, asserting there is exactly one."
|
||||
found = self.points(name, attributes)
|
||||
assert len(found) == 1, (
|
||||
f"expected exactly one {name} point matching {attributes}, "
|
||||
f"got {len(found)}: {found}"
|
||||
)
|
||||
return found[0]
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def otel_metrics():
|
||||
"""
|
||||
Function-scoped metrics collector. Drains any delta state accumulated by
|
||||
earlier tests before yielding, so counts start from zero.
|
||||
"""
|
||||
pytest.importorskip("opentelemetry.sdk")
|
||||
if _otel_metric_reader is None:
|
||||
pytest.skip("OpenTelemetry SDK meter provider was not installed")
|
||||
_otel_metric_reader.get_metrics_data()
|
||||
yield MetricsCollector(_otel_metric_reader)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def bare_ds():
|
||||
"""
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue