Commit graph

16 commits

Author SHA1 Message Date
Alex Garcia
033bd9eb4f Document metric exemplars, which Datasette already emits
Histograms recorded inside a sampled span carry trace IDs automatically, so a
latency spike links to a trace that caused it. Nothing said so.

Two things are documented because they were measured rather than assumed: an
exemplar is kept per histogram bucket, so the bucket boundaries fixed earlier
in this stack took the same workload from one reachable trace to four; and the
pinned opentelemetry-exporter-prometheus drops exemplars entirely, so the path
that works is an OTLP collector rather than Datasette's Prometheus exporter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

(cherry picked from 9d066255; section numbering and cross-references
adjusted to this branch's demo README, and the exemplar reference placed
as a subsection of the new Metric reference in internals.rst.)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2h9ANGZ7paWSpqs5DUAcG
2026-09-01 17:09:35 -07:00
Alex Garcia
8b35d4a80d Add OpenTelemetry metrics for SQL thread pool saturation and query latency
Spans describe requests that have finished. They structurally cannot answer
"am I saturating my 3 SQL threads right now", because that is a level rather
than an event - and with num_sql_threads defaulting to 3, it is usually the
first thing worth knowing about a busy Datasette. This adds the metrics that
answer it.

Five observable gauges, computed only when something is collecting, so an
instance with no MeterProvider installed does no work for them at all:

  datasette.sql.threads.limit         num_sql_threads
  datasette.sql.threads.queue_depth   queries waiting for a free thread
  datasette.sql.queries.pending       in-flight reads, by db.namespace
  datasette.write.queue_depth         writes behind the single write thread
  datasette.connections.open          tracked file connections

Three instruments recorded inline, which matters because metrics survive
trace sampling and spans do not - an operator sampling 1% of traces still
gets 100% of the latency distribution:

  db.client.operation.duration        semconv histogram, with error.type
  datasette.write.queue_wait          the metric twin of the existing span
  datasette.sql.queries.interrupted   sql_time_limit_ms kills

The interrupted counter closes a gap the plan called out as unanswerable:
"how often are we killing queries at the limit" is a rate, and a rate cannot
be recovered from sampled spans.

Core still creates no provider of any kind, so the architecture is unchanged;
`grep -rn 'opentelemetry.sdk' datasette/` stays empty. One real difference
from tracing is worth recording: _ProxyMeter and its instruments forward to a
provider installed after they were created, whereas ProxyTracer permanently
caches the first concrete tracer it resolves. Module-level instruments are
therefore safe and the test fixture has no ordering constraint.

Live instances are tracked in a lock-guarded WeakSet so instrumenting an
instance never keeps it alive. The pool gauges carry no attribute saying
which Datasette produced them: production runs one instance per process, and
adding an id to disambiguate the test suite's hundreds of instances would buy
unbounded attribute cardinality to fix a case that does not occur. The
collision is documented instead, and the gauge callbacks are plain generator
functions so tests can assert exact values by calling them directly rather
than through the SDK's last-value aggregation.

demos/otel/metrics_demo.py fires 12 concurrent 40ms queries at a 3-thread
pool and samples the gauges mid-flight: queue_depth peaks at exactly 9, and
the duration histogram reads max=0.1695s for a query whose work is 40ms. That
gap is the queue, and it is the thing traces alone will not show you.

Also corrects the demo README's privacy section, which still claimed
parameter values are never recorded - that stopped being unconditionally true
when trace_sql_parameters landed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

(cherry picked from 6ef0dd8c and adapted to the rebuilt phase-1 stack:
attribute names now come from telemetry_registry where entries exist, the
meter carries the instrumentation-scope version and schema URL, and the
interrupted-queries counter skips expected timeouts - callers that opted
into a deliberately short budget, like facet suggestion - matching how
those are excluded from span error status. The internals.rst reference
lands with the registry commit that follows.)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2h9ANGZ7paWSpqs5DUAcG
2026-09-01 17:09:35 -07:00
Alex Garcia
76893cb8d3 Apply black to the OTLP demo receiver
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01F2h9ANGZ7paWSpqs5DUAcG
2026-09-01 17:09:35 -07:00
Alex Garcia
4a97e577b8 Add a pure-Python OTLP demo: receiver, Jaeger recipe, no Docker
demos/otel/ shows the whole export path end to end with uv alone. The
~150 line otlp_receiver.py decodes the real OTLP/HTTP protobuf wire
format and prints a summary on Ctrl-C; a Justfile wraps the receiver,
an instrumented `datasette` under opentelemetry-instrument, a
generated 200-row demo database, and Jaeger from its own binary.
Jaeger and the receiver both listen on 4318, so the Datasette side is
one identical `just serve` either way.

Verified against this branch: startup plus one request against the
demo table exports 93 spans with exactly two roots - the request span
and datasette.startup - and in Jaeger the same run lands as two
traces, 67 spans nested under `GET <route>` and 26 under startup.
That "no orphans" shape is new since the earlier demo iteration: the
per-request server span and the startup parent are in core now, so
the old asgi_wrapper plugin and its "ignore the ~24 tiny traces"
caveats are gone rather than ported.

serve_forever() runs on a worker thread with the main thread waiting
on an event, and SIGINT/SIGTERM are handled explicitly, because
KeyboardInterrupt alone does not reliably reach the script through
the `uv run` wrapper - the summary was silently never printed. The
per-batch print flushes explicitly so live feedback survives a pipe.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012U7coQfVu8nK2R4q2mCULA
2026-09-01 16:24:35 -07:00
Simon Willison
5ea7098e4d Fixed an unnecessary f-string 2024-02-04 10:15:21 -08:00
Cameron Yick
452a587e23
JavaScript Plugin API, providing custom panels and column menu items
Thanks, Cameron Yick.

https://github.com/simonw/datasette/pull/2052

Co-authored-by: Simon Willison <swillison@gmail.com>
2023-10-12 17:00:27 -07:00
Simon Willison
9676b2deb0 Upgrade Docker images to Python 3.11, closes #1853 2022-10-25 12:04:53 -07:00
Simon Willison
080d4b3e06 Switch to python:3.10.6-slim-bullseye for datasette publish - refs #1768 2022-08-14 08:49:14 -07:00
Simon Willison
10659c3f1f datasette-debug-asgi plugin to help investigate #1590 2022-01-13 16:38:53 -08:00
Simon Willison
ed77eda6d8 Add datasette-redirect-to-https plugin
Also configured suprvisord children to log to stdout, so that I
can see them with flyctly logs -a datasette-apache-proxy-demo

Refs #1524
2021-11-20 15:30:25 -08:00
Simon Willison
f11a13d73f Extract out Apache config to separate file, refs #1524 2021-11-20 12:23:40 -08:00
Simon Willison
48951e4304 Switch to hosting demo on Fly, closes #1522 2021-11-20 10:51:51 -08:00
Simon Willison
494f11d5cc Switch from Alpine to Debian, refs #1522 2021-11-20 10:51:14 -08:00
Simon Willison
24b5006ad7 ProxyPreserveHost On for apache-proxy demo, refs #1522 2021-11-19 17:11:13 -08:00
Simon Willison
a1ba6cd6bb Use build arguments, refs #1522 2021-11-19 16:34:35 -08:00
Simon Willison
c76bbd4066 New live demo with Apache proxying, refs #1522 2021-11-19 14:50:06 -08:00