From 90736074a00cbc8a074d99eb9b4b78a50f1919cb Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Mon, 21 Sep 2026 13:50:09 +0200 Subject: [PATCH 01/51] rfc(feature): Capture SDK options Add initial boilerplate draft for capturing Sentry SDK init options. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/XXXX-capture-sdk-options.md | 90 ++++++++++++++++++++++++++++++++ 1 file changed, 90 insertions(+) create mode 100644 text/XXXX-capture-sdk-options.md diff --git a/text/XXXX-capture-sdk-options.md b/text/XXXX-capture-sdk-options.md new file mode 100644 index 00000000..3cb4e2bc --- /dev/null +++ b/text/XXXX-capture-sdk-options.md @@ -0,0 +1,90 @@ +- Start Date: 2026-09-21 +- RFC Type: feature +- RFC PR: +- RFC Status: draft +- RFC Author: @mydea +- RFC Approver: + +# Summary + +Today, we have no visibility into which options a given Sentry SDK instance was configured +with. This RFC proposes a mechanism for SDKs to report the configuration they were +initialized with (the arguments passed to `Sentry.init()`, plus relevant derived/effective +values) to Sentry, so that this information can be stored, surfaced, and acted upon. + +Knowing the configured options unlocks a range of use cases — from self-healing and +debugging ("was this simply not enabled, or did the event never get sent?"), to product +analytics (which options are actually used, which matter most), to user-facing features +(showing all `Sentry.init()`s producing data into a project, auditing setups, and eventually +allowing configuration changes from the UI). We do not need to build all of these up front; +this RFC focuses on the capture and transport mechanism, with the downstream use cases +described as motivation and future work. + +# Motivation + +We currently cannot answer basic questions about how an SDK instance is configured. This +creates gaps in several areas: + +- **Self-healing / support**: When data is missing or filtered, we cannot tell whether a + feature was disabled, a sample rate dropped the event, or something failed. Knowing the + configured options would let us distinguish "never enabled" from "enabled but filtered". +- **Analytics & product decisions**: We do not know which options are actually used in the + wild. This data would inform what to highlight in docs, what to deprecate or remove in + major versions, and where to invest. +- **Configuration-aware querying**: Configuration changes can affect trends over time (e.g. + a change to filtering or sampling rules). Surfacing these changes would help explain shifts + in data. +- **Setup audits & warnings**: We could audit `init()` setups to suggest improvements or warn + users about potentially confusing behavior given their settings. +- **Discoverability of data sources**: Users could see all of the `Sentry.init()`s producing + data into a project, understand the sources of their data, and see the filtering, sampling, + and other configuration that affects it — including changes over time. + +Ideas of what to eventually do with this data (not all in scope for this initial project): + +- Analytics of which options are used, for decision-making about docs, deprecations, and majors. +- Flagging configuration changes that might affect trends over time when querying. +- Auditing `init()` setups to suggest helpful changes or warn about confusing behavior. +- Showing users a list of all the `Sentry.init()`s producing data into each project, so they + can understand their data sources and any filtering/sampling/config that affects them. +- Showing changes to configuration over time. +- A UI for turning sources on and off and making configuration changes (e.g. via Seer PRs). +- A clear discoverability element — showing which data sources are configured to produce (and + not produce) which telemetry types. + +# Background + + + +# Supporting Data + + + +# Options Considered + + + +# Drawbacks + + + +# Unresolved questions + +- What is the minimum viable set of options to capture for the initial project? +- What is the serialization format and how do we handle function/integration options? +- How do we prevent leaking secrets or PII contained in configuration? +- What is the transport and storage model, and how do we deduplicate across many clients? +- How does this generalize beyond JS to other SDKs? From ff1dc8cced770b89cc5815ae1b6ae7f05cc0c391 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Mon, 21 Sep 2026 13:54:22 +0200 Subject: [PATCH 02/51] rfc(feature): Add Background, drop Supporting Data Describe that we do not capture SDK config today, and the adjacent signals (event SDK metadata, client reports). Remove the empty Supporting Data section for now. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/XXXX-capture-sdk-options.md | 39 ++++++++++++++++++++++++++------ 1 file changed, 32 insertions(+), 7 deletions(-) diff --git a/text/XXXX-capture-sdk-options.md b/text/XXXX-capture-sdk-options.md index 3cb4e2bc..4abc718e 100644 --- a/text/XXXX-capture-sdk-options.md +++ b/text/XXXX-capture-sdk-options.md @@ -54,15 +54,40 @@ Ideas of what to eventually do with this data (not all in scope for this initial # Background - -# Supporting Data - - - # Options Considered +Broadly, there are two ways to get configuration data from the SDK to Sentry. Both assume +the SDK can produce a serialized view of its options; they differ in _how that view is +transported and handled_. + +## Option A (preferred): A dedicated envelope item for SDK options + +Introduce a new, first-class envelope item type dedicated to SDK configuration. The SDK +serializes its options and sends them as a standalone payload, handled on its own path +server-side rather than being coupled to error/transaction events. + +This is the preferred option because it decouples "what is this instance configured with" +from "did this instance send an event". It gives us a clean, purpose-built payload we can +version, store, and reason about independently, and it is not dependent on an error ever +being produced. It also follows the transport precedent set by client reports (a separate, +non-event envelope item on its own cadence; see Background). + +With this option, the design work is mostly about two questions, which we will dive into in +more detail: + +- **a) The shape of the envelope** — what exactly we put into the payload: which options we + capture, how we represent non-serializable options (functions like `beforeSend`, + integration instances), how we normalize/redact sensitive values, and how we keep the + schema consistent and generalizable across SDKs (this starts with JS). +- **b) How/when to send it, and how/when to store it** — the send cadence (once per init, on + change, periodically, flushed on shutdown), and the server-side handling: where it lands, + how we group instances, how we deduplicate identical configs, and how we track changes over + time. + +## Option B: Expand error events to carry SDK options + +Alternatively, we could piggyback on error (and transaction) events — for example by +expanding the existing `sdk` key on the event to carry the full configured options, and then +doing the work server-side to infer, group, and store this out of the event stream. + +This avoids a new envelope type and reuses an existing, well-understood transport. However, +it inherits the drawbacks of being event-coupled: configuration is only observed when (and +as often as) events are sent, an instance that never produces an event is invisible, +identical config is re-sent on every event (wasteful, needs server-side dedup), and we +overload the event schema and pipeline with data that is not really about the event. The +server-side grouping/storage problem also becomes harder because the signal is buried inside +the high-volume event stream. + +For these reasons Option A is preferred; the remaining sections focus on the questions it +raises. # Drawbacks From be6919992a650bb3feaf126d930a624feb97b478 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Mon, 21 Sep 2026 14:13:43 +0200 Subject: [PATCH 04/51] rfc(feature): Add Envelope Shape section Propose a language-agnostic, primitives-only payload: sdk (identity + opt-in integration status via `active`), meta, options, integration_options, and a free-form `_other` bucket. Cover callback normalization, generic integration options, reflecting which integrations are actually used, and cross-SDK naming. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/XXXX-capture-sdk-options.md | 139 +++++++++++++++++++++++++++++++ 1 file changed, 139 insertions(+) diff --git a/text/XXXX-capture-sdk-options.md b/text/XXXX-capture-sdk-options.md index f1fa933b..8c962def 100644 --- a/text/XXXX-capture-sdk-options.md +++ b/text/XXXX-capture-sdk-options.md @@ -135,6 +135,145 @@ the high-volume event stream. For these reasons Option A is preferred; the remaining sections focus on the questions it raises. +# Envelope Shape + +This section proposes the shape of the dedicated SDK-options payload (question **a** above). +The goal is a **language-agnostic** shape that works equally for JavaScript, Python, and every other SDK, so the server can handle a single, consistent schema. + +## Design principles + +- **Primitives only.** The payload contains only JSON-serializable values: strings, numbers, + booleans, `null`, arrays, and plain objects. No runtime constructs (functions, class + instances, streams, etc.) ever appear literally. +- **Callbacks and other runtime values are reduced to markers.** We do not care about a + callback's implementation, only that _a user-defined callback was set_. Any non-serializable + value is normalized to a sentinel. We reuse the normalization convention SDKs already have + (e.g. JS `normalize()` turning a function into `"[Function: name]"`). So + `beforeSend: (event) => …` becomes `"beforeSend": "[Function]"` (optionally + `"[Function: beforeSend]"`), a class instance becomes `"[SomeType]"`, and so on. +- **Faithful to the init config, and nested.** The `options` block mirrors the structure the + user passed to `init()`, preserving nesting rather than flattening it. +- **Generic representation of integration options.** Integrations are user-configurable + (`Sentry.myIntegration({ filter: 'aaa' })`), so we must capture their options generically — + keyed by integration name, with their options normalized by the same rules. We do not need + to understand any specific integration's options; we just record them. +- **Room for well-known metadata and for open-ended data.** Alongside the raw options there is + space for well-known, structured metadata (SDK identity, release/environment, etc.) and a + free-form bucket SDKs can use for anything not yet modeled. + +## Proposed shape + +```json +{ + "sdk": { + "name": "sentry.javascript.node", + "version": "10.0.0", + "packages": [{ "name": "npm:@sentry/node", "version": "10.0.0" }], + "integrations": { + "InboundFilters": {}, + "ExpressIntegration": { "active": true }, + "FastifyIntegration": { "active": false }, + "KoaIntegration": {}, + "MyIntegration": {} + } + }, + + "meta": { + "release": "my-app@1.2.3", + "environment": "production", + "dist": "42", + "runtime": { "name": "node", "version": "20.11.0" } + }, + + "options": { + "dsn": "https://@o0.ingest.sentry.io/0", + "sampleRate": 1.0, + "tracesSampleRate": 0.2, + "sendDefaultPii": true, + "debug": false, + "beforeSend": "[Function]", + "tracesSampler": "[Function]", + "denyUrls": ["https://example.com/ignore"], + "integrations": ["InboundFilters", "MyIntegration"] + }, + + "integration_options": { + "InboundFilters": {}, + "MyIntegration": { + "filter": "aaa", + "shouldLog": "[Function]" + } + }, + + "_other": {} +} +``` + +## Fields + +- **`sdk`** — SDK identity metadata: `name`, `version`, and `packages`. This is the same + information SDKs attach to error/transaction events today; **this RFC proposes moving it + here** so it lives in one canonical place. It also carries the `integrations` map — the set + of registered integrations plus their opt-in runtime status (see below). The configured + _options_ of each integration live in the separate top-level `integration_options` block; + the `sdk.integrations` map is about integration _identity and status_. +- **`meta`** — well-known, general metadata that we want first-class regardless of how it was + set: `release`, `environment`, `dist`, and runtime/platform information (e.g. runtime name + and version). These describe the instance producing data, complementing the raw `options`. +- **`options`** — a normalized snapshot of the configuration passed to `init()`, nested to + mirror the user's input, with all values reduced to primitives per the rules above. The + `integrations` init option is represented here as a list of names; the details live in the + dedicated `integration_options` block to avoid duplicating (and bloating) the raw options. +- **`integration_options`** — a map of integration name → its normalized options. This is how + we generically capture things like `MyIntegration({ filter: 'aaa' })` without understanding + any specific integration. Integrations with no options serialize to `{}`. +- **`_other`** — a free-form, SDK-defined bucket for anything not covered by the well-known + fields above. The leading underscore signals that this is arbitrary, unstructured data. + Keeps the schema forward-compatible: SDKs can record additional data without a schema change, + and useful keys can later be promoted to first-class fields. + +## Reflecting which integrations are actually used + +Knowing which integrations are _registered_ is not the same as knowing which are actually +_doing anything_. In the Node SDK, for example, a large set of integrations is added by +default (`ExpressIntegration`, `FastifyIntegration`, `KoaIntegration`, …), but a given app +typically uses only one of them. For analytics and audits we care about the difference +between "this integration is present because it ships by default" and "this integration is +actually instrumenting this app". + +To capture this generically without enumerating every integration, `sdk.integrations` is a +map keyed by integration name, where the value is a small, **opt-in** status object: + +- The **presence of a key** means the integration is registered/enabled. +- An optional **`active`** boolean means the integration determined at runtime whether it is + actually in effect. `ExpressIntegration` sets `active: true` once it successfully patches + Express; a defaulted integration whose target framework is absent can report + `active: false`. An integration that reports nothing leaves its value as `{}` (`active` + simply absent / unknown). + +Key properties of this design: + +- **Generic.** No integration-specific fields in the schema; any integration can contribute + the well-known `active` signal (and we can add further opt-in status keys later). +- **Opt-in and non-exhaustive.** Integrations are not required to report status. We selectively + push this into the integrations where the signal is valuable to us (e.g. the framework + integrations), and leave the rest unreported. +- **Distinct from options.** This map carries identity/status only; configured options remain + in the top-level `integration_options` block. + +Note there is a timing implication: `active` is often only known slightly after `init()` (once +instrumentation runs), which influences _when_ the payload is sent or updated. This is +discussed in the send/store section. + +## Cross-SDK naming + +Option keys differ across SDKs (JS `tracesSampleRate` vs. Python `traces_sample_rate`). We +need to decide whether the payload uses each SDK's native option names as-is, or a canonical +cross-SDK naming so the server can compare the same option across languages. A canonical +catalog (in the spirit of [0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md)) +would make analytics far easier but requires each SDK to map its options; native names are +simpler but push normalization to the server. This is called out as an open question. + # Drawbacks +# Drawbacks -# Unresolved questions +- **Risk of leaking sensitive data.** Configuration can contain secrets and PII — DSNs, auth + tokens or custom headers in transport options, tunnel URLs, and user-meaningful values in + options like `denyUrls`, `initialScope`, or `serverName`. Even with primitives-only + normalization, we are shipping user configuration to Sentry, and redaction is imperfect. Any + new field an SDK captures is a potential leak, so the capture surface must be curated + carefully. +- **Ingestion and storage overhead.** This is a brand-new payload type sent at potentially very + high volume (every init for client/serverless traffic). Even with sampling and dedup, it adds + ingestion load, a new storage model, and server-side processing that did not exist before. +- **Data may be incomplete or misleading.** The techniques that keep volume down also reduce + fidelity: client sampling can under-sample or miss low-traffic releases and rare + configurations, and dedup-by-release keeps only the first-seen config per release even when + config actually varies within a release. Decisions made on this data (e.g. deprecating an + option that "looks unused") could be based on a non-representative picture. +- **Lossy representation of runtime options.** Reducing callbacks to `"[Function]"` tells us a + `beforeSend`/`tracesSampler` exists but nothing about what it does. For audit-style use cases + ("warn about confusing behavior") this is a hard limit — we can see that filtering is + configured, not what it filters. +- **Ongoing maintenance burden.** The normalized schema (and any cross-SDK naming catalog) must + be kept in sync with evolving options and integrations across every SDK and language. New + options are invisible until each SDK is updated to capture them, so the dataset always lags + the SDKs, and consistency across SDKs takes continuous effort. +- **Added SDK complexity and runtime cost.** Every SDK gains new machinery: debounced sending, + flush-on-shutdown, opt-in per-integration `active` tracking, normalization, and (for clients) + sampling. This is more code, more surface for bugs, and some runtime overhead on every init. +- **Unresolved dedup identity.** Storage relies on a good key to deduplicate by, and it is not + yet clear that `release` (or any single field) is sufficient (see Storing). Getting this wrong + means either storing too much or collapsing genuinely different configurations together. -- What is the minimum viable set of options to capture for the initial project? -- What is the serialization format and how do we handle function/integration options? -- How do we prevent leaking secrets or PII contained in configuration? -- What is the transport and storage model, and how do we deduplicate across many clients? -- How does this generalize beyond JS to other SDKs? From 32c38fa5c4c003067de0ce5167e024e55614517b Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Mon, 21 Sep 2026 14:30:33 +0200 Subject: [PATCH 06/51] rfc(feature): Assign RFC number 0162, add README index entry Rename XXXX -> 0162 now that the PR is open, fill in the RFC PR link, and add the index entry to README. Co-Authored-By: Claude Opus 4.8 (1M context) --- README.md | 1 + ...{XXXX-capture-sdk-options.md => 0162-capture-sdk-options.md} | 2 +- 2 files changed, 2 insertions(+), 1 deletion(-) rename text/{XXXX-capture-sdk-options.md => 0162-capture-sdk-options.md} (99%) diff --git a/README.md b/README.md index ded193dc..9cc92420 100644 --- a/README.md +++ b/README.md @@ -74,3 +74,4 @@ This repository contains RFCs and DACIs. Lost? - [0152-sdk-symbolicated-frames](text/0152-sdk-symbolicated-frames.md): This RFC proposes a mechanism for SDKs to mark stack frames as already symbolicated on the client side, so that the backend (processing/symbolicator) can skip symbolication for those frames - [0153-decoupling-sentrys-generative-ai-conventions-from-open-telemetry](text/0153-decoupling-sentrys-generative-ai-conventions-from-open-telemetry.md): Decoupling Sentry's Generative AI conventions from OpenTelemetry - [0157-distroless-base-images](text/0157-distroless-base-images.md): Distroless base images +- [0162-capture-sdk-options](text/0162-capture-sdk-options.md): Capture SDK options diff --git a/text/XXXX-capture-sdk-options.md b/text/0162-capture-sdk-options.md similarity index 99% rename from text/XXXX-capture-sdk-options.md rename to text/0162-capture-sdk-options.md index 99000947..db1678d1 100644 --- a/text/XXXX-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -1,6 +1,6 @@ - Start Date: 2026-09-21 - RFC Type: feature -- RFC PR: +- RFC PR: https://github.com/getsentry/rfcs/pull/162 - RFC Status: draft - RFC Author: @mydea - RFC Approver: From c3ffd5f5bd7cc7d064274599f5f5b5b16fab9fd6 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Mon, 21 Sep 2026 14:32:16 +0200 Subject: [PATCH 07/51] rfc(feature): Propose `sdk_config` envelope item type name Add an "Envelope item type" subsection recommending `sdk_config`, with `sdk_options` and `client_config` as considered alternatives. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/0162-capture-sdk-options.md | 11 +++++++++++ 1 file changed, 11 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index db1678d1..9b9b84d3 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -140,6 +140,17 @@ raises. This section proposes the shape of the dedicated SDK-options payload (question **a** above). The goal is a **language-agnostic** shape that works equally for JavaScript, Python, and every other SDK, so the server can handle a single, consistent schema. +## Envelope item type + +We propose naming the new envelope item type **`sdk_config`**, following the existing +snake_case convention for item types (`event`, `transaction`, `client_report`, `session`, …). +`sdk_config` reads well because the payload is broader than just the raw `init()` options — it +also carries SDK identity, integration status, and general metadata. + +Alternatives considered: `sdk_options` (closest to the literal `init()` arguments, but narrower +than what the payload actually contains) and `client_config` (risks confusion with Sentry +"client reports" and with the SDK's internal `Client`). We recommend `sdk_config`. + ## Design principles - **Primitives only.** The payload contains only JSON-serializable values: strings, numbers, From 2bff0afa47ee0ec5802d27a3f509c2b3fe358ab9 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Mon, 21 Sep 2026 14:35:32 +0200 Subject: [PATCH 08/51] rfc(feature): Add follow-up scope; recommend unspecced option names - Add "Not in scope / Follow-up work" (product usage; removing SDK metadata from events, which needs a backfill from sdk_config). - Cross-SDK naming: recommend leaving option names unspecced (native SDK names as-is) rather than a canonical catalog. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/0162-capture-sdk-options.md | 37 ++++++++++++++++++++++++++------ 1 file changed, 31 insertions(+), 6 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 9b9b84d3..b60d0380 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -286,12 +286,21 @@ discussed in the send/store section. ## Cross-SDK naming -Option keys differ across SDKs (JS `tracesSampleRate` vs. Python `traces_sample_rate`). We -need to decide whether the payload uses each SDK's native option names as-is, or a canonical -cross-SDK naming so the server can compare the same option across languages. A canonical -catalog (in the spirit of [0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md)) -would make analytics far easier but requires each SDK to map its options; native names are -simpler but push normalization to the server. This is called out as an open question. +Option keys differ across SDKs (JS `tracesSampleRate` vs. Python `traces_sample_rate`). One +option would be a canonical cross-SDK naming (in the spirit of +[0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md)) so the server can +compare the same option across languages, but that requires every SDK to map its options to the +canonical catalog and keep it in sync. + +**Recommendation: leave the option names unspecced.** By design, the payload uses each SDK's +**native option names as-is** — SDKs simply report options as they are named in that SDK, with +no normalization to a shared vocabulary. This makes cross-SDK analysis harder (the server, or a +consumer, has to reconcile differently-named-but-equivalent options), but it is far easier to +reason about and implement in the SDKs: there is nothing to map, nothing to keep in sync, and +new options are captured automatically without a catalog change. Given the RFC starts with JS +and the primary near-term value is per-SDK/per-release insight, this trade-off is worth it; a +canonical mapping can be layered on later (server-side or in analysis) if cross-SDK comparison +becomes important. # Sending and Storing @@ -442,3 +451,19 @@ settling on release alone. yet clear that `release` (or any single field) is sufficient (see Storing). Getting this wrong means either storing too much or collapsing genuinely different configurations together. +# Not in scope / Follow-up work + +This RFC focuses on capturing, transporting, and storing SDK configuration. The following are +explicitly out of scope here and left as follow-up work: + +- **Actually using this data in product.** The downstream use cases from the Motivation + (analytics, setup audits/warnings, showing data sources, configuration-over-time views, a UI + to change configuration, etc.) are not designed here — they build on top of the data this RFC + makes available. +- **Removing SDK metadata from error/transaction events.** Once `sdk_config` is the canonical + home for SDK identity/integration metadata, it makes sense to stop duplicating it on every + error/transaction event. This is not free, though: that metadata is currently searchable and + used on events, so removing it requires a mechanism to backfill it onto events from the stored + `sdk_config` (so events remain searchable/filterable by SDK, version, integrations, etc.). + Designing that backfill is follow-up work. + From cd11ad7f895ba92de49a6eb5dd7f633eadeb430c Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 23 Sep 2026 11:33:19 +0200 Subject: [PATCH 09/51] remove repetitive section --- text/0162-capture-sdk-options.md | 12 ------------ 1 file changed, 12 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index b60d0380..be3933a9 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -40,18 +40,6 @@ creates gaps in several areas: data into a project, understand the sources of their data, and see the filtering, sampling, and other configuration that affects it — including changes over time. -Ideas of what to eventually do with this data (not all in scope for this initial project): - -- Analytics of which options are used, for decision-making about docs, deprecations, and majors. -- Flagging configuration changes that might affect trends over time when querying. -- Auditing `init()` setups to suggest helpful changes or warn about confusing behavior. -- Showing users a list of all the `Sentry.init()`s producing data into each project, so they - can understand their data sources and any filtering/sampling/config that affects them. -- Showing changes to configuration over time. -- A UI for turning sources on and off and making configuration changes (e.g. via Seer PRs). -- A clear discoverability element — showing which data sources are configured to produce (and - not produce) which telemetry types. - # Background Today we do not capture the configuration of an SDK instance in any meaningful, first-class From 510ee37f92f377e3a09e2460dad0dbb7f0bbf088 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 23 Sep 2026 11:40:00 +0200 Subject: [PATCH 10/51] some adjustments --- text/0162-capture-sdk-options.md | 125 ++++++++++++++++++++++++++----- 1 file changed, 106 insertions(+), 19 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index be3933a9..35c448df 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -146,10 +146,9 @@ than what the payload actually contains) and `client_config` (risks confusion wi instances, streams, etc.) ever appear literally. - **Callbacks and other runtime values are reduced to markers.** We do not care about a callback's implementation, only that _a user-defined callback was set_. Any non-serializable - value is normalized to a sentinel. We reuse the normalization convention SDKs already have - (e.g. JS `normalize()` turning a function into `"[Function: name]"`). So - `beforeSend: (event) => …` becomes `"beforeSend": "[Function]"` (optionally - `"[Function: beforeSend]"`), a class instance becomes `"[SomeType]"`, and so on. + value is normalized to a sentinel, following the exact rules in + [Options serialization rules](#options-serialization-rules) below (functions → `"[Function]"`, + integrations → their name, other runtime constructs → a type marker). - **Faithful to the init config, and nested.** The `options` block mirrors the structure the user passed to `init()`, preserving nesting rather than flattening it. - **Generic representation of integration options.** Integrations are user-configurable @@ -160,8 +159,39 @@ than what the payload actually contains) and `client_config` (risks confusion wi space for well-known, structured metadata (SDK identity, release/environment, etc.) and a free-form bucket SDKs can use for anything not yet modeled. +## Options serialization rules + +When serializing the `options` block (and `integration_options`), SDKs apply the following rules, +top to bottom, to every value. The goal is a deterministic, primitives-only representation that is +consistent across SDKs. + +- **Primitives pass through.** Strings, numbers, booleans, and `null` are emitted as-is. Arrays + and plain objects are emitted structurally, with each element/value serialized by these same + rules (nesting is preserved). +- **Functions become `"[Function]"`.** Any callback (`beforeSend`, `tracesSampler`, + `beforeBreadcrumb`, transport factories, etc.) is replaced by the literal string `"[Function]"`. + SDKs MAY include the function name when readily available (`"[Function: beforeSend]"`), but the + bare `"[Function]"` marker is the required baseline — consumers must not depend on the name. +- **Integrations are replaced by their name.** In the `options.integrations` list, each configured + integration is serialized to its **integration name string** (e.g. `MyIntegration` → + `"MyIntegration"`), never the integration instance/object. The integration's _own_ configured + options are captured separately, keyed by that same name, in the top-level `integration_options` + block. So `Sentry.init({ integrations: [Sentry.myIntegration({ filter: 'aaa' })] })` yields + `"integrations": ["MyIntegration"]` in `options` and + `"MyIntegration": { "filter": "aaa" }` in `integration_options`. +- **Other runtime constructs become a type marker.** Any remaining non-serializable value (a class + instance, stream, socket, etc.) is replaced by a bracketed type marker, e.g. `"[SomeType]"`, + reusing each SDK's existing normalization convention (e.g. JS `normalize()`). + +These rules are what "normalized" means throughout this document. SDKs do **not** scrub sensitive +values as part of serialization — redaction is a separate, server-side concern (see `options`). + ## Proposed shape +The shape below is the **stored** payload. SDKs send everything here **except** +`normalized_options`, which Relay derives at ingestion time from `options` (see +[Normalization at ingestion](#normalization-at-ingestion)). + ```json { "timestamp": "2026-09-21T12:00:00Z", @@ -198,6 +228,14 @@ than what the payload actually contains) and `client_config` (risks confusion wi "integrations": ["InboundFilters", "MyIntegration"] }, + "normalized_options": { + "sample_rate": { "key": "sampleRate", "value": 1.0 }, + "traces_sample_rate": { "key": "tracesSampleRate", "value": 0.2 }, + "send_default_pii": { "key": "sendDefaultPii", "value": true }, + "debug": { "key": "debug", "value": false }, + "before_send": { "key": "beforeSend", "value": "[Function]" } + }, + "integration_options": { "InboundFilters": {}, "MyIntegration": { @@ -231,6 +269,10 @@ than what the payload actually contains) and `client_config` (risks confusion wi dedicated `integration_options` block to avoid duplicating (and bloating) the raw options. Fields in `options` should be **scrubbed server-side** for sensitive data (especially tokens and other secrets); notably, SDKs do **not** perform any client-side scrubbing of these values. +- **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical + cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps + a canonical key to `{ "key": , "value": }`, so consumers + can compare the same option across SDKs while still seeing what it was called natively. - **`integration_options`** — a map of integration name → its normalized options. This is how we generically capture things like `MyIntegration({ filter: 'aaa' })` without understanding any specific integration. Integrations with no options serialize to `{}`. @@ -274,21 +316,66 @@ discussed in the send/store section. ## Cross-SDK naming -Option keys differ across SDKs (JS `tracesSampleRate` vs. Python `traces_sample_rate`). One -option would be a canonical cross-SDK naming (in the spirit of -[0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md)) so the server can -compare the same option across languages, but that requires every SDK to map its options to the -canonical catalog and keep it in sync. - -**Recommendation: leave the option names unspecced.** By design, the payload uses each SDK's -**native option names as-is** — SDKs simply report options as they are named in that SDK, with -no normalization to a shared vocabulary. This makes cross-SDK analysis harder (the server, or a -consumer, has to reconcile differently-named-but-equivalent options), but it is far easier to -reason about and implement in the SDKs: there is nothing to map, nothing to keep in sync, and -new options are captured automatically without a catalog change. Given the RFC starts with JS -and the primary near-term value is per-SDK/per-release insight, this trade-off is worth it; a -canonical mapping can be layered on later (server-side or in analysis) if cross-SDK comparison -becomes important. +Option keys differ across SDKs (JS `tracesSampleRate` vs. Python `traces_sample_rate`). To +compare the same option across languages we want a canonical cross-SDK vocabulary (in the spirit +of [0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md)), but we do **not** +want every SDK to own that mapping and keep it in sync. + +**Recommendation: SDKs send native names; Relay normalizes a defined subset.** The `options` +block keeps each SDK's **native option names as-is** — SDKs report options exactly as they are +named in that SDK, with no normalization to a shared vocabulary. This keeps the SDK side dumb and +maintenance-free: there is nothing to map, nothing to keep in sync, and new options are captured +automatically. The canonical mapping lives **server-side in Relay**, which reads a fixed set of +well-known options out of `options` and emits them into `normalized_options` under canonical keys +(see below). + +This gets us the best of both: full, native fidelity in `options`, plus a normalized, +cross-SDK-comparable view in `normalized_options` — without pushing catalog upkeep into every SDK. + +## Normalization at ingestion + +`normalized_options` is produced by **Relay at ingestion time**, never sent by the SDK. For each +option in the normalization catalog, Relay looks it up in the incoming `options` (by the native +key registered for that SDK) and, if present, emits an entry: + +```json +"normalized_options": { + "traces_sample_rate": { "key": "tracesSampleRate", "value": 0.5 } +} +``` + +- The **outer key** is the canonical, cross-SDK name (snake_case). +- **`key`** is the native option name it was normalized from, so the original naming is not lost. +- **`value`** is the value of this options. + +Only options in the catalog are normalized; everything else remains available under `options`. +Because the mapping is Relay-side, **extending the catalog is a Relay change only** — no SDK +release or rollout is needed to start normalizing a new option, and back-data already stored as +raw `options` can be re-normalized. + +### Which options we normalize + +We start with a small, curated set of high-value options that are meaningful across SDKs. The +canonical key is the same across languages; only the native `key` differs. + +| Canonical key | Meaning | Native examples (JS → Python) | +| ---------------------- | -------------------------------- | ----------------------------------------- | +| `sample_rate` | Error sample rate | `sampleRate` → `sample_rate` | +| `traces_sample_rate` | Tracing sample rate | `tracesSampleRate` → `traces_sample_rate` | +| `profiles_sample_rate` | Profiling sample rate | `profilesSampleRate` → `profiles_sample_rate` | +| `send_default_pii` | Whether default PII is sent | `sendDefaultPii` → `send_default_pii` | +| `debug` | Debug logging enabled | `debug` → `debug` | +| `enabled` | Whether the SDK is enabled | `enabled` → `enabled` | +| `before_send` | Whether a `before_send` hook is set (marker) | `beforeSend` → `before_send` | +| `before_send_transaction` | Whether a `before_send_transaction` hook is set (marker) | `beforeSendTransaction` → `before_send_transaction` | +| `before_send_span` | Whether a `before_send_span` hook is set (marker) | `beforeSendSpan` → `before_send_span` | +| `ignore_spans` | Span-ignore rules | `ignoreSpans` → `ignore_spans` | +| `traces_sampler` | Whether a `traces_sampler` hook is set (marker) | `tracesSampler` → `traces_sampler` | + +`release`, `environment`, and `dist` are deliberately **not** in this catalog — they are already +promoted to first-class fields under `meta`. The catalog is intended to grow over time; the list +above is the initial, deliberately-conservative set, and the exact registry (canonical key ↔ +per-SDK native key) is maintained alongside Relay. # Sending and Storing From ad51f497ff6b2ba1feb5490ec7c8ec8e45a8a08f Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 23 Sep 2026 16:18:56 +0200 Subject: [PATCH 11/51] rfc(feature): Split effective options from user-set keys; normalize in Relay Rework the envelope shape so SDKs send their final/effective `options` as-is plus a flat `options_set_by_user` provenance array, and Relay derives `normalized_options` (a canonical-keyed subset) at ingestion. Add explicit serialization rules (functions -> [Function], integrations -> name), the normalization catalog, and rationale for both decisions (Relay-side normalization; effective options + flat set-by-user array). Co-Authored-By: Claude Opus 4.8 (1M context) --- text/0162-capture-sdk-options.md | 103 +++++++++++++++++++++++++++---- 1 file changed, 90 insertions(+), 13 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 35c448df..90c8c2f1 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -149,8 +149,10 @@ than what the payload actually contains) and `client_config` (risks confusion wi value is normalized to a sentinel, following the exact rules in [Options serialization rules](#options-serialization-rules) below (functions → `"[Function]"`, integrations → their name, other runtime constructs → a type marker). -- **Faithful to the init config, and nested.** The `options` block mirrors the structure the - user passed to `init()`, preserving nesting rather than flattening it. +- **Effective config, natively shaped and nested.** The `options` block carries the SDK's final, + effective options (after defaults and derivation), keyed by native option names and preserving + the nesting the user would recognize rather than flattening it. Which of those keys the user + explicitly set is recorded separately in `options_set_by_user`. - **Generic representation of integration options.** Integrations are user-configurable (`Sentry.myIntegration({ filter: 'aaa' })`), so we must capture their options generically — keyed by integration name, with their options normalized by the same rules. We do not need @@ -222,12 +224,22 @@ The shape below is the **stored** payload. SDKs send everything here **except** "tracesSampleRate": 0.2, "sendDefaultPii": true, "debug": false, + "environment": "production", "beforeSend": "[Function]", "tracesSampler": "[Function]", "denyUrls": ["https://example.com/ignore"], "integrations": ["InboundFilters", "MyIntegration"] }, + "options_set_by_user": [ + "dsn", + "tracesSampleRate", + "sendDefaultPii", + "beforeSend", + "tracesSampler", + "denyUrls" + ], + "normalized_options": { "sample_rate": { "key": "sampleRate", "value": 1.0 }, "traces_sample_rate": { "key": "tracesSampleRate", "value": 0.2 }, @@ -263,16 +275,32 @@ The shape below is the **stored** payload. SDKs send everything here **except** - **`meta`** — well-known, general metadata that we want first-class regardless of how it was set: `release`, `environment`, `dist`, and runtime/platform information (e.g. runtime name and version). These describe the instance producing data, complementing the raw `options`. -- **`options`** — a normalized snapshot of the configuration passed to `init()`, nested to - mirror the user's input, with all values reduced to primitives per the rules above. The - `integrations` init option is represented here as a list of names; the details live in the - dedicated `integration_options` block to avoid duplicating (and bloating) the raw options. - Fields in `options` should be **scrubbed server-side** for sensitive data (especially tokens - and other secrets); notably, SDKs do **not** perform any client-side scrubbing of these values. +- **`options`** — a normalized snapshot of the SDK's **final, effective** configuration (i.e. the + options object the SDK actually runs with, after defaults, env-var resolution, and any + derived/integration-injected values), nested to mirror the user's input, with all values reduced + to primitives per the rules above. Sending the effective config — rather than only the literal + `init()` arguments — is what lets us answer behavior questions ("what sample rate is actually in + effect", "is it enabled"); it also reads directly off the SDK's existing options object with no + extra plumbing, and reflects the settled state at send time (see the debounce in Sending). The + `integrations` option is represented here as a list of names; the details live in the dedicated + `integration_options` block to avoid duplicating (and bloating) the raw options. Fields in + `options` should be **scrubbed server-side** for sensitive data (especially tokens and other + secrets); notably, SDKs do **not** perform any client-side scrubbing of these values. +- **`options_set_by_user`** — a flat array of the **native option keys the user explicitly set** + in `init()` (as opposed to values that came from defaults, env vars, or integrations). This is + the signal that lets us distinguish "the user chose this" from "this is just a + default" — essential for adoption analytics and setup audits, where a defaulted value is not + "usage". The SDK produces it by diffing the keys of the raw `init()` argument against the + effective `options`. It lists **top-level native keys only** (no nested paths); if nested + provenance is ever needed it can be added later. Keys here always correspond to keys present in + `options`. This array is the **single, uniform source of provenance**: to ask "did the user set + option X?" for _any_ option (normalized or not), check whether its native key is a member — + `normalized_options` deliberately does **not** duplicate this signal. - **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps a canonical key to `{ "key": , "value": }`, so consumers - can compare the same option across SDKs while still seeing what it was called natively. + can compare the same option across SDKs while still seeing what it was called natively. Whether + the user set it is answered the same way as for any other option — via `options_set_by_user`. - **`integration_options`** — a map of integration name → its normalized options. This is how we generically capture things like `MyIntegration({ filter: 'aaa' })` without understanding any specific integration. Integrations with no options serialize to `{}`. @@ -281,6 +309,33 @@ The shape below is the **stored** payload. SDKs send everything here **except** Keeps the schema forward-compatible: SDKs can record additional data without a schema change, and useful keys can later be promoted to first-class fields. +## Why effective options plus a flat set-by-user array + +We deliberately split configuration into two SDK-sent pieces — the full **effective `options`** +and a flat **`options_set_by_user`** array — rather than, say, shipping both a user-provided and an +effective options tree, or wrapping every option value in a `{ value, source }` object. The +benefits: + +- **Easy to implement in SDKs.** `options` is essentially the SDK's existing effective options + object, serialized — most SDKs can read it straight from an existing API (e.g. JS + `client.getOptions()`, or the equivalent effective-options accessor in other SDKs), so there is + no second config snapshot to capture or keep around. `options_set_by_user` is produced by a + single diff of the raw `init()` argument's keys against that object. No per-option plumbing, no + wrapper types, no bookkeeping threaded through the option system. +- **Easy to reason about.** `options` means exactly one thing — the configuration the SDK actually + runs with — and `options_set_by_user` means exactly one thing — which of those the user chose. + There is no ambiguity about whether a given block is "before" or "after" defaults, and no mixed + value/metadata shape to interpret. What each field represents is obvious from its name. +- **Can be joined as needed.** Keeping provenance as a separate flat set means any consumer can + answer "was this user-set?" for _any_ option with a simple membership check, and can just as + easily ignore provenance entirely when it does not care. The two pieces compose on demand + (including for `normalized_options`, which joins against the same array) instead of being + pre-fused into one heavier structure that every consumer pays for whether or not they need it. + +The trade-off is that provenance is not co-located with each value (you look it up rather than +reading it inline), and the array is top-level-only. Both are acceptable given how much simpler +this keeps the SDK side and the schema. + ## Reflecting which integrations are actually used Knowing which integrations are _registered_ is not the same as knowing which are actually @@ -332,6 +387,26 @@ well-known options out of `options` and emits them into `normalized_options` und This gets us the best of both: full, native fidelity in `options`, plus a normalized, cross-SDK-comparable view in `normalized_options` — without pushing catalog upkeep into every SDK. +**Why Relay, not the SDK.** Doing the normalization server-side is a deliberate choice, for +several reinforcing reasons: + +- **Simpler SDKs.** Each SDK only has to serialize and send its own native options (something it + effectively already has). It does not need to know the canonical vocabulary, map its keys onto + it, or reason about equivalence across languages — that logic never ships in the SDK at all. +- **Nothing to keep aligned across SDKs.** A canonical mapping owned by the SDKs would have to be + implemented, and kept consistent, in _every_ SDK and language independently. Any drift (a + mismatched canonical key, a missed option, an inconsistent value normalization) would silently + corrupt cross-SDK comparisons. Centralizing removes that entire class of cross-SDK + synchronization problem. +- **One place to implement and maintain.** The normalization logic and the knowledge of which + canonical keys exist live in a **single system (Relay)** rather than being duplicated across N + SDKs. There is one implementation to write, test, review, and reason about — not one per SDK. +- **Changeable over time without SDK releases.** What we choose to normalize will evolve. Because + the catalog lives in Relay, we can add, rename, or refine normalized keys — and re-normalize + already-ingested `options` — as a Relay change alone, with **no SDK update, release, or user + upgrade** required. If normalization lived in the SDKs, every change would mean shipping every + SDK and waiting for the ecosystem to upgrade, so the normalized dataset would always lag. + ## Normalization at ingestion `normalized_options` is produced by **Relay at ingestion time**, never sent by the SDK. For each @@ -346,12 +421,14 @@ key registered for that SDK) and, if present, emits an entry: - The **outer key** is the canonical, cross-SDK name (snake_case). - **`key`** is the native option name it was normalized from, so the original naming is not lost. -- **`value`** is the value of this options. +- **`value`** is the value of this option (the effective value from `options`). + +Provenance is intentionally not repeated here: to check whether a normalized option was user-set, +look up its native `key` in `options_set_by_user`, exactly as you would for any raw option. Only options in the catalog are normalized; everything else remains available under `options`. -Because the mapping is Relay-side, **extending the catalog is a Relay change only** — no SDK -release or rollout is needed to start normalizing a new option, and back-data already stored as -raw `options` can be re-normalized. +Because the mapping is Relay-side (see [Why Relay, not the SDK](#cross-sdk-naming)), the catalog +can be extended without an SDK release, and already-stored raw `options` can be re-normalized. ### Which options we normalize From 7df85fe2215b4ff2528bd1123f34ec0dee5cff6c Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 23 Sep 2026 16:20:45 +0200 Subject: [PATCH 12/51] adjustments --- text/0162-capture-sdk-options.md | 14 ++++++-------- 1 file changed, 6 insertions(+), 8 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 90c8c2f1..20944945 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -288,14 +288,12 @@ The shape below is the **stored** payload. SDKs send everything here **except** secrets); notably, SDKs do **not** perform any client-side scrubbing of these values. - **`options_set_by_user`** — a flat array of the **native option keys the user explicitly set** in `init()` (as opposed to values that came from defaults, env vars, or integrations). This is - the signal that lets us distinguish "the user chose this" from "this is just a - default" — essential for adoption analytics and setup audits, where a defaulted value is not - "usage". The SDK produces it by diffing the keys of the raw `init()` argument against the - effective `options`. It lists **top-level native keys only** (no nested paths); if nested - provenance is ever needed it can be added later. Keys here always correspond to keys present in - `options`. This array is the **single, uniform source of provenance**: to ask "did the user set - option X?" for _any_ option (normalized or not), check whether its native key is a member — - `normalized_options` deliberately does **not** duplicate this signal. + the signal that lets us distinguish default values from user-set values + — essential for adoption analytics and setup audits, where a default value is not + "usage". This is effectively similar to `Object.keys(options)` where `options` are the user-provided options for `Sentry.init(options)`. + It lists **top-level native keys only** (no nested paths); if nested + keys are ever needed they can be added later. Keys here always correspond to keys present in + `options`. - **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps a canonical key to `{ "key": , "value": }`, so consumers From ccc43f1916113b3f133839d5d9613c8976d2a1f8 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 23 Sep 2026 16:47:04 +0200 Subject: [PATCH 13/51] rfc(feature): Rename integration status to `applied`; flexible send timing; add Open Questions MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Rename the per-integration runtime signal `active` -> `applied` (clearer: "took effect at runtime" rather than enabled/disabled). - Reframe server-SDK send timing: debounce is one example default; any settling hook/mechanism is fine as long as final options are captured. - Note the new `sdk_config` envelope item degrades gracefully — older Relay/self-hosted discards the unknown item and processes the rest. - Add an Open Questions section (new-envelope-type infra: rate limits, payload/size limits, data category/quota/billing, outcomes) and reference the release dedup-key question. - Clarify scrubbing: primarily server-side; SDKs MAY scrub known-sensitive values as best-effort, without guaranteeing fully-scrubbed data. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/0162-capture-sdk-options.md | 103 ++++++++++++++++++++++--------- 1 file changed, 74 insertions(+), 29 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 20944945..5a86c64e 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -139,6 +139,15 @@ Alternatives considered: `sdk_options` (closest to the literal `init()` argument than what the payload actually contains) and `client_config` (risks confusion with Sentry "client reports" and with the SDK's internal `Client`). We recommend `sdk_config`. +**Backward compatibility with older ingest / self-hosted.** Introducing a new envelope item type +is safe for older infrastructure. Envelopes are designed so that unknown item types are simply +**ignored**: an older Relay / self-hosted Sentry that predates `sdk_config` will not recognize the +item and will **discard it**, while still processing the other items in the same envelope (errors, +transactions, etc.) as usual. That is the desired behavior here — SDKs can start emitting +`sdk_config` unconditionally, and setups that cannot yet handle it lose only this +new-and-supplementary payload, with no impact on existing data. No SDK-side version gating is +required. + ## Design principles - **Primitives only.** The payload contains only JSON-serializable values: strings, numbers, @@ -185,8 +194,9 @@ consistent across SDKs. instance, stream, socket, etc.) is replaced by a bracketed type marker, e.g. `"[SomeType]"`, reusing each SDK's existing normalization convention (e.g. JS `normalize()`). -These rules are what "normalized" means throughout this document. SDKs do **not** scrub sensitive -values as part of serialization — redaction is a separate, server-side concern (see `options`). +These rules are what "normalized" means throughout this document. Scrubbing sensitive values is a +**separate, primarily server-side concern** (see `options`) — SDKs **MAY** scrub values they know +to be sensitive, but they do not attempt to guarantee fully-scrubbed data. ## Proposed shape @@ -204,8 +214,8 @@ The shape below is the **stored** payload. SDKs send everything here **except** "packages": [{ "name": "npm:@sentry/node", "version": "10.0.0" }], "integrations": { "InboundFilters": {}, - "ExpressIntegration": { "active": true }, - "FastifyIntegration": { "active": false }, + "ExpressIntegration": { "applied": true }, + "FastifyIntegration": { "applied": false }, "KoaIntegration": {}, "MyIntegration": {} } @@ -283,9 +293,11 @@ The shape below is the **stored** payload. SDKs send everything here **except** effect", "is it enabled"); it also reads directly off the SDK's existing options object with no extra plumbing, and reflects the settled state at send time (see the debounce in Sending). The `integrations` option is represented here as a list of names; the details live in the dedicated - `integration_options` block to avoid duplicating (and bloating) the raw options. Fields in - `options` should be **scrubbed server-side** for sensitive data (especially tokens and other - secrets); notably, SDKs do **not** perform any client-side scrubbing of these values. + `integration_options` block to avoid duplicating (and bloating) the raw options. Sensitive data + (especially tokens and other secrets) is **primarily scrubbed server-side**; SDKs **MAY** + additionally scrub values they know to be sensitive (e.g. a field that always holds a secret), + but this is a best-effort defense-in-depth measure — SDKs do **not** attempt to guarantee + fully-scrubbed data, and the server-side scrubbing remains the mechanism we rely on. - **`options_set_by_user`** — a flat array of the **native option keys the user explicitly set** in `init()` (as opposed to values that came from defaults, env vars, or integrations). This is the signal that lets us distinguish default values from user-set values @@ -347,23 +359,25 @@ To capture this generically without enumerating every integration, `sdk.integrat map keyed by integration name, where the value is a small, **opt-in** status object: - The **presence of a key** means the integration is registered/enabled. -- An optional **`active`** boolean means the integration determined at runtime whether it is - actually in effect. `ExpressIntegration` sets `active: true` once it successfully patches +- An optional **`applied`** boolean means the integration determined at runtime whether it + actually took effect. `ExpressIntegration` sets `applied: true` once it successfully patches Express; a defaulted integration whose target framework is absent can report - `active: false`. An integration that reports nothing leaves its value as `{}` (`active` - simply absent / unknown). + `applied: false`. An integration that reports nothing leaves its value as `{}` (`applied` + simply absent / unknown). We chose `applied` over `active` because "active" reads as + enabled/disabled — which the key's mere presence already conveys — whereas `applied` names the + distinct signal we want: the integration ran and took effect at runtime. Key properties of this design: - **Generic.** No integration-specific fields in the schema; any integration can contribute - the well-known `active` signal (and we can add further opt-in status keys later). + the well-known `applied` signal (and we can add further opt-in status keys later). - **Opt-in and non-exhaustive.** Integrations are not required to report status. We selectively push this into the integrations where the signal is valuable to us (e.g. the framework integrations), and leave the rest unreported. - **Distinct from options.** This map carries identity/status only; configured options remain in the top-level `integration_options` block. -Note there is a timing implication: `active` is often only known slightly after `init()` (once +Note there is a timing implication: `applied` is often only known slightly after `init()` (once instrumentation runs), which influences _when_ the payload is sent or updated. This is discussed in the send/store section. @@ -462,26 +476,35 @@ We start with the server case; the client case is trickier and is covered next. ## Sending — Server SDKs -For server SDKs, the payload is sent **once, shortly after `init()`**, guarded by a short -**debounce delay**: - -- **Debounced send after init.** Rather than sending synchronously at the end of `init()`, the - SDK schedules the send after a short delay (default on the order of **2 seconds**). The delay - lets the configuration **stabilize**: some values are not known at the exact moment `init()` - returns — for example an integration's runtime `active` status (see "Reflecting which - integrations are actually used"), lazily-registered integrations, or release/environment - detected asynchronously. Debouncing collapses this settling period into a single payload that - reflects the effective configuration rather than a half-initialized snapshot. -- **Configurable wait period.** The delay **MAY be configurable**. Each SDK should pick a - sensible default based on **when the data it cares about becomes available** — an SDK that - only detects framework instrumentation after the first request may want a longer default than - one whose config is fully known almost immediately. Exposing it as an option lets specific - setups tune it. +For server SDKs, the payload is sent **once, shortly after `init()`**, but only once the +configuration has **settled**. The hard requirement is only this: + +- **Send the final, settled options — not a half-initialized snapshot.** Some values are not + known at the exact moment `init()` returns — for example an integration's runtime `applied` + status (see "Reflecting which integrations are actually used"), lazily-registered integrations, + or release/environment detected asynchronously. The SDK must wait until it can reasonably expect + these to have settled before capturing and sending the payload. - **One payload per init.** The goal is a single, stable payload per SDK instance/init, not a stream of updates. If the config meaningfully changes later, that is handled as a separate concern (and mostly matters for long-lived processes); the common case is one send per process start. +**How to wait is up to each SDK.** We deliberately do not mandate a single mechanism — each SDK +**MAY** use whatever hook or mechanism fits its runtime and lifecycle, as long as it reasonably +captures the final options. A few examples: + +- **A short debounce delay** (a reasonable default, e.g. on the order of **2 seconds**): schedule + the send a short time after `init()` rather than synchronously, letting the settling period + collapse into one payload. The delay **MAY be configurable** so setups that settle slower (e.g. + an SDK that only detects framework instrumentation after the first request) can tune it. +- **A lifecycle hook** the SDK already has — e.g. sending after the first request/transaction is + processed, on an "SDK ready"/post-init hook, or when the event loop first goes idle. +- **Any equivalent trigger** that reliably fires after the config the SDK cares about is known. + +Debouncing is the simplest baseline and a fine default, but it is an example, not a rule: an SDK +with a natural hook that guarantees settled options should prefer that. What matters is the +outcome — the payload reflects the effective configuration. + **Short-lived processes.** Some server environments (serverless functions, short CLI invocations) may exit before the debounce timer fires. In those cases the SDK should flush the pending options payload on shutdown / at the same points it already flushes events, so the @@ -570,6 +593,28 @@ in some setups. We may need a composite key (e.g. release + environment, or a ha normalized options) or a different identifier altogether. This needs to be validated before settling on release alone. +# Open Questions + +- **What does standing up a new envelope item type actually require?** Introducing `sdk_config` + is not just an SDK + storage change; it needs first-class handling in the ingest pipeline, and + we need to enumerate what that entails before committing. At least: + - **Relay / ingest support.** Registering the new item type, routing it to its own handler, + and any validation/normalization (the `normalized_options` derivation lands here). + - **Rate limiting.** Does `sdk_config` need its own rate-limit category, or does it share an + existing one? How does rate limiting interact with the send cadence (debounced server sends, + client sampling)? What is communicated back to the SDK (e.g. `429` / `Retry-After`) and how + does the SDK back off? + - **Payload / size limits.** What per-item and per-envelope size limits apply, and what happens + when a config exceeds them — reject, or truncate (and if so, how, given nested `options`)? + - **Data category, quota & billing.** Which data category does it map to for + quotas/outcomes/billing? The intent is that this is supplementary telemetry, so it most + likely should **not** be billed like events — but that needs to be decided explicitly. + - **Outcomes / observability.** How are dropped or rejected `sdk_config` items recorded + (outcomes, reasons) so we can see ingestion health for this new type? +- **Is `release` the right dedup key?** (See [Storing](#storing) above.) Whether release alone + identifies a distinct configuration, or we need a composite key / options hash, still needs + validation. + # Drawbacks - **Risk of leaking sensitive data.** Configuration can contain secrets and PII — DSNs, auth @@ -595,7 +640,7 @@ settling on release alone. options are invisible until each SDK is updated to capture them, so the dataset always lags the SDKs, and consistency across SDKs takes continuous effort. - **Added SDK complexity and runtime cost.** Every SDK gains new machinery: debounced sending, - flush-on-shutdown, opt-in per-integration `active` tracking, normalization, and (for clients) + flush-on-shutdown, opt-in per-integration `applied` tracking, normalization, and (for clients) sampling. This is more code, more surface for bugs, and some runtime overhead on every init. - **Unresolved dedup identity.** Storage relies on a good key to deduplicate by, and it is not yet clear that `release` (or any single field) is sufficient (see Storing). Getting this wrong From 04d7e45e4dc6f41bf17d479f0eb870e363efb253 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 24 Sep 2026 09:49:27 +0200 Subject: [PATCH 14/51] rfc(feature): Consolidate integration data into one `integrations` map Replace the two parallel name-keyed maps (`sdk.integrations` for status, top-level `integration_options` for options) with a single top-level `integrations` map. Each entry holds the integration's runtime `applied` status plus its configured options nested under an `options` key, so arbitrary user options can never collide with well-known status keys. Also document a known limitation of sending effective options: the user-provided *form* is lost when the SDK coerces a value (notably `integrations` passed as a function -> resolved array). Note this is rare (essentially `integrations` and `stackParser` in JS) and can be layered on later via a sparse `options_user_provided` block without reworking the shape. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/0162-capture-sdk-options.md | 104 +++++++++++++++++++------------ 1 file changed, 65 insertions(+), 39 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 5a86c64e..e9033f5a 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -172,9 +172,9 @@ required. ## Options serialization rules -When serializing the `options` block (and `integration_options`), SDKs apply the following rules, -top to bottom, to every value. The goal is a deterministic, primitives-only representation that is -consistent across SDKs. +When serializing the `options` block (and each integration's `options`), SDKs apply the following +rules, top to bottom, to every value. The goal is a deterministic, primitives-only representation +that is consistent across SDKs. - **Primitives pass through.** Strings, numbers, booleans, and `null` are emitted as-is. Arrays and plain objects are emitted structurally, with each element/value serialized by these same @@ -186,10 +186,10 @@ consistent across SDKs. - **Integrations are replaced by their name.** In the `options.integrations` list, each configured integration is serialized to its **integration name string** (e.g. `MyIntegration` → `"MyIntegration"`), never the integration instance/object. The integration's _own_ configured - options are captured separately, keyed by that same name, in the top-level `integration_options` - block. So `Sentry.init({ integrations: [Sentry.myIntegration({ filter: 'aaa' })] })` yields - `"integrations": ["MyIntegration"]` in `options` and - `"MyIntegration": { "filter": "aaa" }` in `integration_options`. + options are captured separately, keyed by that same name, in the top-level `integrations` block. + So `Sentry.init({ integrations: [Sentry.myIntegration({ filter: 'aaa' })] })` yields + `"integrations": ["MyIntegration"]` in `options` and, in the `integrations` block, + `"MyIntegration": { "options": { "filter": "aaa" } }`. - **Other runtime constructs become a type marker.** Any remaining non-serializable value (a class instance, stream, socket, etc.) is replaced by a bracketed type marker, e.g. `"[SomeType]"`, reusing each SDK's existing normalization convention (e.g. JS `normalize()`). @@ -211,13 +211,16 @@ The shape below is the **stored** payload. SDKs send everything here **except** "sdk": { "name": "sentry.javascript.node", "version": "10.0.0", - "packages": [{ "name": "npm:@sentry/node", "version": "10.0.0" }], - "integrations": { - "InboundFilters": {}, - "ExpressIntegration": { "applied": true }, - "FastifyIntegration": { "applied": false }, - "KoaIntegration": {}, - "MyIntegration": {} + "packages": [{ "name": "npm:@sentry/node", "version": "10.0.0" }] + }, + + "integrations": { + "InboundFilters": { "options": {} }, + "ExpressIntegration": { "applied": true, "options": {} }, + "FastifyIntegration": { "applied": false, "options": {} }, + "KoaIntegration": { "options": {} }, + "MyIntegration": { + "options": { "filter": "aaa", "shouldLog": "[Function]" } } }, @@ -258,14 +261,6 @@ The shape below is the **stored** payload. SDKs send everything here **except** "before_send": { "key": "beforeSend", "value": "[Function]" } }, - "integration_options": { - "InboundFilters": {}, - "MyIntegration": { - "filter": "aaa", - "shouldLog": "[Function]" - } - }, - "_other": {} } ``` @@ -278,10 +273,19 @@ The shape below is the **stored** payload. SDKs send everything here **except** convention (ISO 8601 shown here; could equally be epoch seconds to match event `timestamp`). - **`sdk`** — SDK identity metadata: `name`, `version`, and `packages`. This is the same information SDKs attach to error/transaction events today; **this RFC proposes moving it - here** so it lives in one canonical place. It also carries the `integrations` map — the set - of registered integrations plus their opt-in runtime status (see below). The configured - _options_ of each integration live in the separate top-level `integration_options` block; - the `sdk.integrations` map is about integration _identity and status_. + here** so it lives in one canonical place. (The set of integrations, previously part of the + event's `sdk` object as a list of names, is captured more richly in the top-level + `integrations` block below.) +- **`integrations`** — a map keyed by integration name, where each entry holds everything we know + about that integration in one place: its runtime status and its configured options. This + consolidates what would otherwise be two parallel name-keyed maps. Per entry: + - The **presence of the key** means the integration is registered/enabled. + - An optional, opt-in **`applied`** boolean records whether the integration determined at runtime + that it actually took effect (see "Reflecting which integrations are actually used"). + - **`options`** is the integration's own normalized options, nested under this key so arbitrary + user options can never collide with well-known status keys like `applied`. Integrations with no + options report `"options": {}`. This is how we generically capture things like + `MyIntegration({ filter: 'aaa' })` without understanding any specific integration. - **`meta`** — well-known, general metadata that we want first-class regardless of how it was set: `release`, `environment`, `dist`, and runtime/platform information (e.g. runtime name and version). These describe the instance producing data, complementing the raw `options`. @@ -292,8 +296,9 @@ The shape below is the **stored** payload. SDKs send everything here **except** `init()` arguments — is what lets us answer behavior questions ("what sample rate is actually in effect", "is it enabled"); it also reads directly off the SDK's existing options object with no extra plumbing, and reflects the settled state at send time (see the debounce in Sending). The - `integrations` option is represented here as a list of names; the details live in the dedicated - `integration_options` block to avoid duplicating (and bloating) the raw options. Sensitive data + `integrations` option is represented here as a list of names; each integration's identity, + runtime status, and configured options live in the dedicated top-level `integrations` block, to + avoid duplicating (and bloating) the raw options. Sensitive data (especially tokens and other secrets) is **primarily scrubbed server-side**; SDKs **MAY** additionally scrub values they know to be sensitive (e.g. a field that always holds a secret), but this is a best-effort defense-in-depth measure — SDKs do **not** attempt to guarantee @@ -311,9 +316,6 @@ The shape below is the **stored** payload. SDKs send everything here **except** a canonical key to `{ "key": , "value": }`, so consumers can compare the same option across SDKs while still seeing what it was called natively. Whether the user set it is answered the same way as for any other option — via `options_set_by_user`. -- **`integration_options`** — a map of integration name → its normalized options. This is how - we generically capture things like `MyIntegration({ filter: 'aaa' })` without understanding - any specific integration. Integrations with no options serialize to `{}`. - **`_other`** — a free-form, SDK-defined bucket for anything not covered by the well-known fields above. The leading underscore signals that this is arbitrary, unstructured data. Keeps the schema forward-compatible: SDKs can record additional data without a schema change, @@ -346,6 +348,32 @@ The trade-off is that provenance is not co-located with each value (you look it reading it inline), and the array is top-level-only. Both are acceptable given how much simpler this keeps the SDK side and the schema. +### Known limitation: user-provided _form_ is not preserved + +Because we send the **effective** options, we lose information when the user expressed a value in a +form that the SDK coerces into something else before it lands on the effective options. The value +we report is the coerced result, and `options_set_by_user` only tells us the option _was_ set — not +_how_ it was written. Concretely: a user can pass `integrations` as either an array or a +**function** (`(defaults) => Integration[]`), but the client always ends up holding a resolved +array — so we cannot tell, from the payload, whether the user configured integrations via a +function. We therefore cannot answer questions like "how many users pass `integrations` as a +function". + +We accept this limitation for now: + +- **It is rare.** In the JS SDK, the option type system pins genuine user-vs-effective _shape_ + divergence to essentially two options: `integrations` (array-or-function → array) and + `stackParser` (array-or-function → function). Everything else keeps the same shape; only defaults + or env values get filled in, which the effective `options` already captures faithfully. +- **`integrations` is the main case that matters**, and even there the question ("was it a + function?") is a nice-to-have, not core to the primary use cases. + +If we later decide this signal is worth capturing, we can layer it on **without reworking the +shape** — e.g. an optional, sparse `options_user_provided` block that records the user-provided +(serialized) value _only_ for the few keys whose provided form differs from the effective one (so +it would carry `{ "integrations": "[Function]" }` and otherwise be empty). We deliberately leave +that out of the initial design and revisit it only if a concrete need arises. + ## Reflecting which integrations are actually used Knowing which integrations are _registered_ is not the same as knowing which are actually @@ -355,17 +383,14 @@ typically uses only one of them. For analytics and audits we care about the diff between "this integration is present because it ships by default" and "this integration is actually instrumenting this app". -To capture this generically without enumerating every integration, `sdk.integrations` is a -map keyed by integration name, where the value is a small, **opt-in** status object: +This is captured generically, without enumerating any specific integration, via the `applied` key +on each entry of the top-level `integrations` map: - The **presence of a key** means the integration is registered/enabled. - An optional **`applied`** boolean means the integration determined at runtime whether it actually took effect. `ExpressIntegration` sets `applied: true` once it successfully patches Express; a defaulted integration whose target framework is absent can report - `applied: false`. An integration that reports nothing leaves its value as `{}` (`applied` - simply absent / unknown). We chose `applied` over `active` because "active" reads as - enabled/disabled — which the key's mere presence already conveys — whereas `applied` names the - distinct signal we want: the integration ran and took effect at runtime. + `applied: false`. An integration that reports nothing simply omits `applied` (absent / unknown). Key properties of this design: @@ -374,8 +399,9 @@ Key properties of this design: - **Opt-in and non-exhaustive.** Integrations are not required to report status. We selectively push this into the integrations where the signal is valuable to us (e.g. the framework integrations), and leave the rest unreported. -- **Distinct from options.** This map carries identity/status only; configured options remain - in the top-level `integration_options` block. +- **Namespaced from options.** `applied` (and any future status key) sits at the top of the entry, + while the integration's own configured options are nested under the entry's `options` key, so + the two never collide. Note there is a timing implication: `applied` is often only known slightly after `init()` (once instrumentation runs), which influences _when_ the payload is sent or updated. This is From fcf03d08e41f9a0399c34523e481e19eff9b9faf Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 24 Sep 2026 09:57:11 +0200 Subject: [PATCH 15/51] reorder --- text/0162-capture-sdk-options.md | 20 ++++++++++---------- 1 file changed, 10 insertions(+), 10 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index e9033f5a..20e79109 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -214,16 +214,6 @@ The shape below is the **stored** payload. SDKs send everything here **except** "packages": [{ "name": "npm:@sentry/node", "version": "10.0.0" }] }, - "integrations": { - "InboundFilters": { "options": {} }, - "ExpressIntegration": { "applied": true, "options": {} }, - "FastifyIntegration": { "applied": false, "options": {} }, - "KoaIntegration": { "options": {} }, - "MyIntegration": { - "options": { "filter": "aaa", "shouldLog": "[Function]" } - } - }, - "meta": { "release": "my-app@1.2.3", "environment": "production", @@ -261,6 +251,16 @@ The shape below is the **stored** payload. SDKs send everything here **except** "before_send": { "key": "beforeSend", "value": "[Function]" } }, + "integrations": { + "InboundFilters": { "options": {} }, + "ExpressIntegration": { "applied": true, "options": {} }, + "FastifyIntegration": { "applied": false, "options": {} }, + "KoaIntegration": { "options": {} }, + "MyIntegration": { + "options": { "filter": "aaa", "shouldLog": "[Function]" } + } + }, + "_other": {} } ``` From d3bcd6914d22c3ec853ae873c44de328743e8be8 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 24 Sep 2026 10:57:44 +0200 Subject: [PATCH 16/51] rfc(feature): Unify dedup into composite key + optional SDK-set options hash MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Replace the two-strategy + hybrid dedup framing with a single rule: deduplicate by `release` + `environment` + `dist` plus an optional, SDK-provided options hash that is stable across instances sharing a config. When present, the server folds the hash into the dedup key (capturing within-release variation); when absent, dedup falls back to the composite key. The hash is SDK-side by design (a robust server-side canonical hash isn't feasible) and, if used, MUST be stamped on every event — as an attribute on spans/logs and a context field on error/transaction events — so events correlate exactly to their config. Correlation now falls out of the key: bucket-level without a hash, exact with one. Update Open Questions and the Drawbacks bullets to match. Co-Authored-By: Claude Opus 4.8 (1M context) --- text/0162-capture-sdk-options.md | 115 +++++++++++++++++++++++-------- 1 file changed, 85 insertions(+), 30 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 20e79109..d5987e7a 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -597,27 +597,78 @@ Persist every payload we receive as its own record. where near-identical payloads arrive constantly. The vast majority of records are duplicates that add no information. -### Option 2 (recommended): Deduplicate and store a single record - -Store only **one** record per configuration and discard the rest. Because configuration is -essentially constant per release, we need a key to deduplicate by, and we suggest **release**: - -- We store the **first** payload seen for a given release. -- All subsequent payloads for that **same** release are **discarded** (they are expected to be - identical). -- When a payload arrives with a **new/different release**, we store that as the new record for - that release. - -This keeps storage bounded (roughly one record per release), matches the reality that config -changes track releases, and naturally gives us a per-release history of configuration over time. - -**Open question — is `release` the right dedup key?** We need to verify that release is -sufficient to identify a distinct configuration. It may not always be: config can differ -_within_ the same release (e.g. environment-dependent options, feature flags, per-deployment -overrides, or code paths that call `init()` with different options), and release may be unset -in some setups. We may need a composite key (e.g. release + environment, or a hash of the -normalized options) or a different identifier altogether. This needs to be validated before -settling on release alone. +### Option 2 (recommended): Deduplicate and store one record per configuration + +Store one record per distinct configuration and discard the rest. The vast majority of incoming +payloads are duplicates, so dedup is what keeps this feature affordable. The question is what counts +as a "distinct configuration" — i.e. what we deduplicate by. We propose a **single dedup key with an +optional component**: + +> **`release` + `environment` + `dist` + (optional) SDK-set options hash.** + +#### The composite natural key (`release` + `environment` + `dist`) + +This is the always-available baseline. Release alone is **not** sufficient — configuration can +legitimately differ across environments and builds within the same release (per-environment +options, per-deployment overrides, env-var-driven values) — so the minimum key is +`release` + `environment` + `dist`. + +- It is **bounded and predictable** (roughly one record per tuple), **human-readable and directly + queryable** ("show the config for release X in production"), **cheap** (a few stable top-level + fields, no whole-payload processing), and every event already carries these fields. +- Its weakness is `release`: `environment` effectively always has a value (it defaults to + `production`), `dist` is normally absent, so `release` is the load-bearing part — and it has **no + default** and is frequently unset. When it is, the key degrades to roughly `environment` alone + (almost always `production`), collapsing distinct configs into one bucket. It is also blind to + differences _within_ a tuple: two instances sharing release+environment+dist but differing in some + option are stored as one (first-seen) record, so genuine variation is silently lost. + +The optional hash exists to address exactly that last weakness. + +#### The optional SDK-set options hash + +SDKs **MAY** additionally set a **hash of their options that is stable across instances sharing the +same configuration** (same config → same hash; the hash changes when the config changes). When a +payload carries this hash, **the server includes it in the dedup key**; when it is absent, dedup +falls back to the composite key alone. + +Including the hash means two instances with the same release+environment+dist but genuinely +different options (different hash) are stored as **distinct records** — so within-release variation +and drift become visible instead of being collapsed into the first-seen config. + +- **We recommend hashing the (normalized) options** to produce this value, but SDKs **MAY** choose a + different hashing strategy if it makes more sense for them. The only hard requirement is + stability: the hash must be **identical across instances that share a configuration** and + **differ when the configuration differs**. It does not need to be comparable _across_ SDKs. +- **The hash is computed SDK-side, by design.** Otherwise, we cannot reliable relate events to their respective config. +- **If a hash is used, it MUST also be attached to every event the SDK produces**, so events can be + correlated back to the exact config that produced them: + - as an **attribute** on spans, logs, and other attribute-carrying items. We propose `sentry.config_hash` as a semantic attribute. + - as a **context field** on error and transaction events. We propose `sdk_config.hash` as a new context with a single field for now. + +#### Correlating an event to its config + +This falls directly out of the dedup key: + +- **Without a hash:** correlate via the composite key the event already carries + (`release` + `environment` + `dist`). This is free and needs no event changes, but is only + **bucket-level** — if config varied within the tuple, the event resolves to the bucket (a single + first-seen record, or the set of stored variants), not necessarily the exact config that produced + it. +- **With a hash:** correlation is **exact** — the event's stamped hash matches exactly one stored + `sdk_config` record. This is the whole reason the hash must also live on events. + +#### Trade-offs and caveats of the hash + +- **Per-event overhead.** Stamping the hash adds a field to every event, at full event volume. +- **Timing.** The config is not fully settled the instant `init()` returns (debounce window, + `applied` known late). The SDK must ensure the hash it stamps on events matches the config it + reports — e.g. by hashing only config that is stable from `init()`, or by not stamping until the + config has settled. Events emitted before that point may carry no hash (falling back to + bucket-level correlation). +- **Sampling gaps.** For client/serverless SDKs we _may_ sample `sdk_config` (see Sending), so an + event's hash can reference a config that was **never stored** — correlation then fails for exactly + the instances we chose not to persist. # Open Questions @@ -637,9 +688,10 @@ settling on release alone. likely should **not** be billed like events — but that needs to be decided explicitly. - **Outcomes / observability.** How are dropped or rejected `sdk_config` items recorded (outcomes, reasons) so we can see ingestion health for this new type? -- **Is `release` the right dedup key?** (See [Storing](#storing) above.) Whether release alone - identifies a distinct configuration, or we need a composite key / options hash, still needs - validation. +- **Finalizing the dedup key.** (See [Storing](#storing) above.) We recommend + `release` + `environment` + `dist` plus an optional SDK-set options hash, but the exact field set, + the fallback for the common release-less case, and what SDKs should hash (and how they keep it + stable) still need to be nailed down. # Drawbacks @@ -654,9 +706,10 @@ settling on release alone. ingestion load, a new storage model, and server-side processing that did not exist before. - **Data may be incomplete or misleading.** The techniques that keep volume down also reduce fidelity: client sampling can under-sample or miss low-traffic releases and rare - configurations, and dedup-by-release keeps only the first-seen config per release even when - config actually varies within a release. Decisions made on this data (e.g. deprecating an - option that "looks unused") could be based on a non-representative picture. + configurations, and when no options hash is set, dedup keeps only the first-seen config per + `release`+`environment`+`dist` bucket even when config actually varies within it. Decisions made + on this data (e.g. deprecating an option that "looks unused") could be based on a + non-representative picture. - **Lossy representation of runtime options.** Reducing callbacks to `"[Function]"` tells us a `beforeSend`/`tracesSampler` exists but nothing about what it does. For audit-style use cases ("warn about confusing behavior") this is a hard limit — we can see that filtering is @@ -668,9 +721,11 @@ settling on release alone. - **Added SDK complexity and runtime cost.** Every SDK gains new machinery: debounced sending, flush-on-shutdown, opt-in per-integration `applied` tracking, normalization, and (for clients) sampling. This is more code, more surface for bugs, and some runtime overhead on every init. -- **Unresolved dedup identity.** Storage relies on a good key to deduplicate by, and it is not - yet clear that `release` (or any single field) is sufficient (see Storing). Getting this wrong - means either storing too much or collapsing genuinely different configurations together. +- **Unresolved dedup identity.** Storage relies on a good key to deduplicate by. The composite key + (`release`+`environment`+`dist`) is coarse when `release` is absent, and the finer-grained + optional options hash is only as good as each SDK's hashing (and is not always present). Getting + this wrong means either storing too much or collapsing genuinely different configurations + together (see Storing). # Not in scope / Follow-up work From c215f2f10a3409075d99fea10fd526bbe212cc5b Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 24 Sep 2026 12:56:27 +0200 Subject: [PATCH 17/51] flattening --- text/0162-capture-sdk-options.md | 40 +++++++++++++++++++++----------- 1 file changed, 27 insertions(+), 13 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index d5987e7a..371b5706 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -158,10 +158,11 @@ required. value is normalized to a sentinel, following the exact rules in [Options serialization rules](#options-serialization-rules) below (functions → `"[Function]"`, integrations → their name, other runtime constructs → a type marker). -- **Effective config, natively shaped and nested.** The `options` block carries the SDK's final, - effective options (after defaults and derivation), keyed by native option names and preserving - the nesting the user would recognize rather than flattening it. Which of those keys the user - explicitly set is recorded separately in `options_set_by_user`. +- **Effective config, natively named but flat.** The `options` block carries the SDK's final, + effective options (after defaults and derivation), keyed by native option names. Nested objects + are **flattened with dot notation** (e.g. `dataCollection.http.bodies`) rather than kept as nested + objects, so the structure stays flat and uniform (see the serialization rules). Which of those + keys the user explicitly set is recorded separately in `options_set_by_user`. - **Generic representation of integration options.** Integrations are user-configurable (`Sentry.myIntegration({ filter: 'aaa' })`), so we must capture their options generically — keyed by integration name, with their options normalized by the same rules. We do not need @@ -176,9 +177,14 @@ When serializing the `options` block (and each integration's `options`), SDKs ap rules, top to bottom, to every value. The goal is a deterministic, primitives-only representation that is consistent across SDKs. -- **Primitives pass through.** Strings, numbers, booleans, and `null` are emitted as-is. Arrays - and plain objects are emitted structurally, with each element/value serialized by these same - rules (nesting is preserved). +- **Primitives pass through.** Strings, numbers, booleans, and `null` are emitted as-is. Arrays are + emitted as arrays, with each element serialized by these same rules. +- **Nested objects are flattened with dot notation.** We deliberately keep the structure **flat**: + a nested option object is not emitted as a nested object but as dotted keys joining the path, e.g. + `dataCollection: { http: { bodies: true } }` becomes `"dataCollection.http.bodies": true`. The + same applies inside each integration's `options`. This gives a single, flat, uniform key space + that is easy to store, query, and normalize; the leaf values are serialized by the other rules + here. (Arrays are not flattened — they are kept as arrays.) - **Functions become `"[Function]"`.** Any callback (`beforeSend`, `tracesSampler`, `beforeBreadcrumb`, transport factories, etc.) is replaced by the literal string `"[Function]"`. SDKs MAY include the function name when readily available (`"[Function: beforeSend]"`), but the @@ -231,6 +237,8 @@ The shape below is the **stored** payload. SDKs send everything here **except** "beforeSend": "[Function]", "tracesSampler": "[Function]", "denyUrls": ["https://example.com/ignore"], + "dataCollection.http.bodies": true, + "dataCollection.http.headers": false, "integrations": ["InboundFilters", "MyIntegration"] }, @@ -291,8 +299,8 @@ The shape below is the **stored** payload. SDKs send everything here **except** and version). These describe the instance producing data, complementing the raw `options`. - **`options`** — a normalized snapshot of the SDK's **final, effective** configuration (i.e. the options object the SDK actually runs with, after defaults, env-var resolution, and any - derived/integration-injected values), nested to mirror the user's input, with all values reduced - to primitives per the rules above. Sending the effective config — rather than only the literal + derived/integration-injected values), keyed by native names and flattened with dot notation, with + all values reduced to primitives per the rules above. Sending the effective config — rather than only the literal `init()` arguments — is what lets us answer behavior questions ("what sample rate is actually in effect", "is it enabled"); it also reads directly off the SDK's existing options object with no extra plumbing, and reflects the settled state at send time (see the debounce in Sending). The @@ -308,9 +316,9 @@ The shape below is the **stored** payload. SDKs send everything here **except** the signal that lets us distinguish default values from user-set values — essential for adoption analytics and setup audits, where a default value is not "usage". This is effectively similar to `Object.keys(options)` where `options` are the user-provided options for `Sentry.init(options)`. - It lists **top-level native keys only** (no nested paths); if nested - keys are ever needed they can be added later. Keys here always correspond to keys present in - `options`. + Keys here use the **same flattened dot-notation as `options`**, so a user-set nested value is + listed by its dotted leaf key (e.g. `dataCollection.http.bodies`) and always corresponds 1:1 to a + key present in `options`. - **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps a canonical key to `{ "key": , "value": }`, so consumers @@ -486,6 +494,11 @@ canonical key is the same across languages; only the native `key` differs. | `before_send_span` | Whether a `before_send_span` hook is set (marker) | `beforeSendSpan` → `before_send_span` | | `ignore_spans` | Span-ignore rules | `ignoreSpans` → `ignore_spans` | | `traces_sampler` | Whether a `traces_sampler` hook is set (marker) | `tracesSampler` → `traces_sampler` | +| `data_collection.*` | Data-collection settings (all flattened sub-keys) | `dataCollection.*` → `data_collection.*` | + +A trailing `.*` (e.g. `data_collection.*`) denotes a **family of flattened sub-keys**: every dotted +key under that path (`dataCollection.http.bodies`, `dataCollection.http.headers`, …) is normalized, +each becoming its own `normalized_options` entry under the canonical dotted key. `release`, `environment`, and `dist` are deliberately **not** in this catalog — they are already promoted to first-class fields under `meta`. The catalog is intended to grow over time; the list @@ -682,7 +695,8 @@ This falls directly out of the dedup key: client sampling)? What is communicated back to the SDK (e.g. `429` / `Retry-After`) and how does the SDK back off? - **Payload / size limits.** What per-item and per-envelope size limits apply, and what happens - when a config exceeds them — reject, or truncate (and if so, how, given nested `options`)? + when a config exceeds them — reject, or truncate (and if so, which flattened `options` keys to + drop)? - **Data category, quota & billing.** Which data category does it map to for quotas/outcomes/billing? The intent is that this is supplementary telemetry, so it most likely should **not** be billed like events — but that needs to be decided explicitly. From e8631ab9791c614f9c417866a93e23863cca7e1e Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 24 Sep 2026 13:10:12 +0200 Subject: [PATCH 18/51] add hash --- text/0162-capture-sdk-options.md | 19 +++++++++++++++---- 1 file changed, 15 insertions(+), 4 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 371b5706..4d664301 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -251,6 +251,8 @@ The shape below is the **stored** payload. SDKs send everything here **except** "denyUrls" ], + "options_hash": "9f2c1a7e", + "normalized_options": { "sample_rate": { "key": "sampleRate", "value": 1.0 }, "traces_sample_rate": { "key": "tracesSampleRate", "value": 0.2 }, @@ -319,6 +321,15 @@ The shape below is the **stored** payload. SDKs send everything here **except** Keys here use the **same flattened dot-notation as `options`**, so a user-set nested value is listed by its dotted leaf key (e.g. `dataCollection.http.bodies`) and always corresponds 1:1 to a key present in `options`. +- **`options_hash`** (optional) — an SDK-computed hash of the options that is **stable across + instances sharing the same configuration** (same config → same hash). It is the optional + component of the dedup key: when present, the server folds it in so genuinely different configs + within the same `release`+`environment`+`dist` are stored as distinct records; when absent, dedup + falls back to the composite key alone. If an SDK sets it, the **same value must also be stamped on + every event** (attribute on spans/logs, context field on errors/transactions) so events correlate + exactly to their config. See [Storing](#storing) for the full rules and caveats. We recommend + hashing the normalized options, but SDKs may use another stable strategy; it need not be + comparable across SDKs. - **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps a canonical key to `{ "key": , "value": }`, so consumers @@ -640,10 +651,10 @@ The optional hash exists to address exactly that last weakness. #### The optional SDK-set options hash -SDKs **MAY** additionally set a **hash of their options that is stable across instances sharing the -same configuration** (same config → same hash; the hash changes when the config changes). When a -payload carries this hash, **the server includes it in the dedup key**; when it is absent, dedup -falls back to the composite key alone. +SDKs **MAY** additionally set the `options_hash` field on the `sdk_config` payload — a **hash of +their options that is stable across instances sharing the same configuration** (same config → same +hash; the hash changes when the config changes). When a payload carries this hash, **the server +includes it in the dedup key**; when it is absent, dedup falls back to the composite key alone. Including the hash means two instances with the same release+environment+dist but genuinely different options (different hash) are stored as **distinct records** — so within-release variation From 85198f9c3f6d91e71a041bd192e471dd6c971c54 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 08:59:08 +0200 Subject: [PATCH 19/51] clarify hash --- text/0162-capture-sdk-options.md | 15 ++++++++------- 1 file changed, 8 insertions(+), 7 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 4d664301..2f9c9536 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -327,9 +327,9 @@ The shape below is the **stored** payload. SDKs send everything here **except** within the same `release`+`environment`+`dist` are stored as distinct records; when absent, dedup falls back to the composite key alone. If an SDK sets it, the **same value must also be stamped on every event** (attribute on spans/logs, context field on errors/transactions) so events correlate - exactly to their config. See [Storing](#storing) for the full rules and caveats. We recommend - hashing the normalized options, but SDKs may use another stable strategy; it need not be - comparable across SDKs. + exactly to their config. See [Storing](#storing) for the full rules and caveats. It is computed + off the normalized `options` block (the serialized options as defined in the serialization rules), + which makes it deterministic and stable across instances; it need not be comparable across SDKs. - **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps a canonical key to `{ "key": , "value": }`, so consumers @@ -660,10 +660,11 @@ Including the hash means two instances with the same release+environment+dist bu different options (different hash) are stored as **distinct records** — so within-release variation and drift become visible instead of being collapsed into the first-seen config. -- **We recommend hashing the (normalized) options** to produce this value, but SDKs **MAY** choose a - different hashing strategy if it makes more sense for them. The only hard requirement is - stability: the hash must be **identical across instances that share a configuration** and - **differ when the configuration differs**. It does not need to be comparable _across_ SDKs. +- **The hash is computed off the normalized `options`** — the serialized options block as defined + in the [serialization rules](#options-serialization-rules) — which makes it deterministic and + stable: it is **identical across instances that share a configuration** and **differs when the + configuration differs**. It does not need to be comparable _across_ SDKs, so the specific hash + algorithm is up to each SDK. - **The hash is computed SDK-side, by design.** Otherwise, we cannot reliable relate events to their respective config. - **If a hash is used, it MUST also be attached to every event the SDK produces**, so events can be correlated back to the exact config that produced them: From 52e4465671317c9ba348f316ebde658fde2e1ea8 Mon Sep 17 00:00:00 2001 From: Francesco Gringl-Novy Date: Fri, 25 Sep 2026 12:30:23 +0200 Subject: [PATCH 20/51] Update text/0162-capture-sdk-options.md Co-authored-by: Lorenzo Cian <17258265+lcian@users.noreply.github.com> --- text/0162-capture-sdk-options.md | 1 - 1 file changed, 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 2f9c9536..11c07211 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -499,7 +499,6 @@ canonical key is the same across languages; only the native `key` differs. | `profiles_sample_rate` | Profiling sample rate | `profilesSampleRate` → `profiles_sample_rate` | | `send_default_pii` | Whether default PII is sent | `sendDefaultPii` → `send_default_pii` | | `debug` | Debug logging enabled | `debug` → `debug` | -| `enabled` | Whether the SDK is enabled | `enabled` → `enabled` | | `before_send` | Whether a `before_send` hook is set (marker) | `beforeSend` → `before_send` | | `before_send_transaction` | Whether a `before_send_transaction` hook is set (marker) | `beforeSendTransaction` → `before_send_transaction` | | `before_send_span` | Whether a `before_send_span` hook is set (marker) | `beforeSendSpan` → `before_send_span` | From ed3affb86a8bc99703d0e48333498c5062583530 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 12:34:00 +0200 Subject: [PATCH 21/51] add before_send_log --- text/0162-capture-sdk-options.md | 1 + 1 file changed, 1 insertion(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 11c07211..c98bb699 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -502,6 +502,7 @@ canonical key is the same across languages; only the native `key` differs. | `before_send` | Whether a `before_send` hook is set (marker) | `beforeSend` → `before_send` | | `before_send_transaction` | Whether a `before_send_transaction` hook is set (marker) | `beforeSendTransaction` → `before_send_transaction` | | `before_send_span` | Whether a `before_send_span` hook is set (marker) | `beforeSendSpan` → `before_send_span` | +| `before_send_log` | Whether a `before_send_log` hook is set (marker) | `beforeSendLog` → `before_send_log` | | `ignore_spans` | Span-ignore rules | `ignoreSpans` → `ignore_spans` | | `traces_sampler` | Whether a `traces_sampler` hook is set (marker) | `tracesSampler` → `traces_sampler` | | `data_collection.*` | Data-collection settings (all flattened sub-keys) | `dataCollection.*` → `data_collection.*` | From 46441bb07140c8c0878851c523e75c6b4a319096 Mon Sep 17 00:00:00 2001 From: Francesco Gringl-Novy Date: Fri, 25 Sep 2026 12:34:21 +0200 Subject: [PATCH 22/51] Update text/0162-capture-sdk-options.md Co-authored-by: Lorenzo Cian <17258265+lcian@users.noreply.github.com> --- text/0162-capture-sdk-options.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index c98bb699..f3c087d0 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -660,7 +660,7 @@ Including the hash means two instances with the same release+environment+dist bu different options (different hash) are stored as **distinct records** — so within-release variation and drift become visible instead of being collapsed into the first-seen config. -- **The hash is computed off the normalized `options`** — the serialized options block as defined +- **The hash is computed off the serialized options block as defined in the [serialization rules](#options-serialization-rules) — which makes it deterministic and stable: it is **identical across instances that share a configuration** and **differs when the configuration differs**. It does not need to be comparable _across_ SDKs, so the specific hash From 6a197168423db97a72752048acb7fa7c459af699 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 12:35:05 +0200 Subject: [PATCH 23/51] rephrase --- text/0162-capture-sdk-options.md | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index f3c087d0..a6341353 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -151,8 +151,7 @@ required. ## Design principles - **Primitives only.** The payload contains only JSON-serializable values: strings, numbers, - booleans, `null`, arrays, and plain objects. No runtime constructs (functions, class - instances, streams, etc.) ever appear literally. + booleans, `null`, arrays, and plain objects. - **Callbacks and other runtime values are reduced to markers.** We do not care about a callback's implementation, only that _a user-defined callback was set_. Any non-serializable value is normalized to a sentinel, following the exact rules in From f58c91ead78a7ba51f4d5829969472c1b898fd31 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 12:38:24 +0200 Subject: [PATCH 24/51] remove duplicate integrations --- text/0162-capture-sdk-options.md | 31 ++++++++++++++----------------- 1 file changed, 14 insertions(+), 17 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index a6341353..f4cb1e58 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -156,7 +156,7 @@ required. callback's implementation, only that _a user-defined callback was set_. Any non-serializable value is normalized to a sentinel, following the exact rules in [Options serialization rules](#options-serialization-rules) below (functions → `"[Function]"`, - integrations → their name, other runtime constructs → a type marker). + other runtime constructs → a type marker). - **Effective config, natively named but flat.** The `options` block carries the SDK's final, effective options (after defaults and derivation), keyed by native option names. Nested objects are **flattened with dot notation** (e.g. `dataCollection.http.bodies`) rather than kept as nested @@ -188,13 +188,13 @@ that is consistent across SDKs. `beforeBreadcrumb`, transport factories, etc.) is replaced by the literal string `"[Function]"`. SDKs MAY include the function name when readily available (`"[Function: beforeSend]"`), but the bare `"[Function]"` marker is the required baseline — consumers must not depend on the name. -- **Integrations are replaced by their name.** In the `options.integrations` list, each configured - integration is serialized to its **integration name string** (e.g. `MyIntegration` → - `"MyIntegration"`), never the integration instance/object. The integration's _own_ configured - options are captured separately, keyed by that same name, in the top-level `integrations` block. - So `Sentry.init({ integrations: [Sentry.myIntegration({ filter: 'aaa' })] })` yields - `"integrations": ["MyIntegration"]` in `options` and, in the `integrations` block, - `"MyIntegration": { "options": { "filter": "aaa" } }`. +- **Integrations are omitted from `options` entirely.** The `integrations` option is not emitted in + the `options` block at all — everything we capture about integrations (their identity, runtime + status, and own configured options) lives in the dedicated top-level `integrations` block, keyed by + integration name. So `Sentry.init({ integrations: [Sentry.myIntegration({ filter: 'aaa' })] })` + yields no `integrations` key in `options`, and in the `integrations` block + `"MyIntegration": { "options": { "filter": "aaa" } }`. This avoids duplicating the integration + data in two places. - **Other runtime constructs become a type marker.** Any remaining non-serializable value (a class instance, stream, socket, etc.) is replaced by a bracketed type marker, e.g. `"[SomeType]"`, reusing each SDK's existing normalization convention (e.g. JS `normalize()`). @@ -237,8 +237,7 @@ The shape below is the **stored** payload. SDKs send everything here **except** "tracesSampler": "[Function]", "denyUrls": ["https://example.com/ignore"], "dataCollection.http.bodies": true, - "dataCollection.http.headers": false, - "integrations": ["InboundFilters", "MyIntegration"] + "dataCollection.http.headers": false }, "options_set_by_user": [ @@ -304,10 +303,8 @@ The shape below is the **stored** payload. SDKs send everything here **except** all values reduced to primitives per the rules above. Sending the effective config — rather than only the literal `init()` arguments — is what lets us answer behavior questions ("what sample rate is actually in effect", "is it enabled"); it also reads directly off the SDK's existing options object with no - extra plumbing, and reflects the settled state at send time (see the debounce in Sending). The - `integrations` option is represented here as a list of names; each integration's identity, - runtime status, and configured options live in the dedicated top-level `integrations` block, to - avoid duplicating (and bloating) the raw options. Sensitive data + extra plumbing, and reflects the settled state at send time (see the debounce in Sending). `integrations` are omitted here. + Sensitive data (especially tokens and other secrets) is **primarily scrubbed server-side**; SDKs **MAY** additionally scrub values they know to be sensitive (e.g. a field that always holds a secret), but this is a best-effort defense-in-depth measure — SDKs do **not** attempt to guarantee @@ -373,9 +370,9 @@ form that the SDK coerces into something else before it lands on the effective o we report is the coerced result, and `options_set_by_user` only tells us the option _was_ set — not _how_ it was written. Concretely: a user can pass `integrations` as either an array or a **function** (`(defaults) => Integration[]`), but the client always ends up holding a resolved -array — so we cannot tell, from the payload, whether the user configured integrations via a -function. We therefore cannot answer questions like "how many users pass `integrations` as a -function". +array — which is what the top-level `integrations` block reflects — so we cannot tell, from the +payload, whether the user configured integrations via a function. We therefore cannot answer +questions like "how many users pass `integrations` as a function". We accept this limitation for now: From 4f2d5f6f0a1359c776b5a5a894e30bf5db7aa303 Mon Sep 17 00:00:00 2001 From: Francesco Gringl-Novy Date: Fri, 25 Sep 2026 12:47:18 +0200 Subject: [PATCH 25/51] Update text/0162-capture-sdk-options.md Co-authored-by: David Herberth --- text/0162-capture-sdk-options.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index f4cb1e58..1bce837a 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -7,7 +7,7 @@ # Summary -Today, we have no visibility into which options a given Sentry SDK instance was configured +Today, we have very little visibility into which options a given Sentry SDK instance was configured with. This RFC proposes a mechanism for SDKs to report the configuration they were initialized with (the arguments passed to `Sentry.init()`, plus relevant derived/effective values) to Sentry, so that this information can be stored, surfaced, and acted upon. From acabb0c299d9feb9399fbab10b864f8e3f4b5727 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 12:47:12 +0200 Subject: [PATCH 26/51] adjust hash docs --- text/0162-capture-sdk-options.md | 122 ++++++++++++++++--------------- 1 file changed, 64 insertions(+), 58 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 1bce837a..a3c18352 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -317,15 +317,17 @@ The shape below is the **stored** payload. SDKs send everything here **except** Keys here use the **same flattened dot-notation as `options`**, so a user-set nested value is listed by its dotted leaf key (e.g. `dataCollection.http.bodies`) and always corresponds 1:1 to a key present in `options`. -- **`options_hash`** (optional) — an SDK-computed hash of the options that is **stable across - instances sharing the same configuration** (same config → same hash). It is the optional - component of the dedup key: when present, the server folds it in so genuinely different configs - within the same `release`+`environment`+`dist` are stored as distinct records; when absent, dedup - falls back to the composite key alone. If an SDK sets it, the **same value must also be stamped on - every event** (attribute on spans/logs, context field on errors/transactions) so events correlate - exactly to their config. See [Storing](#storing) for the full rules and caveats. It is computed - off the normalized `options` block (the serialized options as defined in the serialization rules), - which makes it deterministic and stable across instances; it need not be comparable across SDKs. +- **`options_hash`** — an SDK-computed hash of the options that is **stable across + instances sharing the same configuration** (same config → same hash). SDKs **MUST** compute and + send it, and **MUST** stamp the **same value on every event** (attribute on spans/logs, context + field on errors/transactions) so events correlate exactly to their config. It is the fine-grained + component of the dedup key: the server folds it in so genuinely different configs within the same + `release`+`environment`+`dist` are stored as distinct records. See [Storing](#storing) for the + full rules and caveats, including the correlation fallback when no matching stored config exists. + It is computed off the serialized `options` & `integrations` blocks (both normalized + by the same [serialization rules](#options-serialization-rules)), so a change to either — including + an integration's own options or its `applied` status — produces a different hash. This makes it + deterministic and stable across instances; it does not need to be comparable across SDKs. - **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps a canonical key to `{ "key": , "value": }`, so consumers @@ -621,16 +623,43 @@ Persist every payload we receive as its own record. Store one record per distinct configuration and discard the rest. The vast majority of incoming payloads are duplicates, so dedup is what keeps this feature affordable. The question is what counts -as a "distinct configuration" — i.e. what we deduplicate by. We propose a **single dedup key with an -optional component**: +as a "distinct configuration" — i.e. what we deduplicate by. We propose a **dedup key built from the +SDK-set options hash, with the composite natural key as a fallback**: -> **`release` + `environment` + `dist` + (optional) SDK-set options hash.** +> **SDK-set options hash (primary) + `release` + `environment` + `dist` (fallback).** -#### The composite natural key (`release` + `environment` + `dist`) +#### The required SDK-set options hash -This is the always-available baseline. Release alone is **not** sufficient — configuration can -legitimately differ across environments and builds within the same release (per-environment -options, per-deployment overrides, env-var-driven values) — so the minimum key is +SDKs **MUST** set the `options_hash` field on the `sdk_config` payload — a **hash of their options +that is stable across instances sharing the same configuration** (same config → same hash; the hash +changes when the config changes). The server uses it as the primary dedup key, so two instances with +the same release+environment+dist but genuinely different options (different hash) are stored as +**distinct records** — within-release variation and drift are visible instead of being collapsed +into the first-seen config. + +- **The hash is computed off both the serialized `options` block and the `integrations` block**, + each normalized by the same [serialization rules](#options-serialization-rules). Since integrations + no longer live in `options`, they must be folded into the hash explicitly and in the same way as + the options — so a change to which integrations are registered, to an integration's own options, or + to its `applied` status yields a different hash. This makes it deterministic and + stable: it is **identical across instances that share a configuration** and **differs when the + configuration differs**. It does not need to be comparable _across_ SDKs, so the specific hash + algorithm is up to each SDK. +- **The hash is computed SDK-side, by design.** Otherwise, we cannot reliably relate events to their + respective config. +- **The hash MUST also be attached to every event the SDK produces**, so events can be + correlated back to the exact config that produced them: + - as an **attribute** on spans, logs, and other attribute-carrying items. We propose `sentry.config_hash` as a semantic attribute. + - as a **context field** on error and transaction events. We propose `sdk_config.hash` as a new context with a single field for now. + +#### The composite natural key (`release` + `environment` + `dist`) — fallback + +Although the hash is required on the wire, a matching stored `sdk_config` record is **not +guaranteed** to exist for a given event's hash (see the caveats below — the config may have been +dropped, sampled out, or stamped before it settled). For those cases we keep the composite natural +key as a fallback lookup. Release alone is **not** sufficient — configuration can legitimately +differ across environments and builds within the same release (per-environment options, +per-deployment overrides, env-var-driven values) — so the fallback key is `release` + `environment` + `dist`. - It is **bounded and predictable** (roughly one record per tuple), **human-readable and directly @@ -640,44 +669,22 @@ options, per-deployment overrides, env-var-driven values) — so the minimum key `production`), `dist` is normally absent, so `release` is the load-bearing part — and it has **no default** and is frequently unset. When it is, the key degrades to roughly `environment` alone (almost always `production`), collapsing distinct configs into one bucket. It is also blind to - differences _within_ a tuple: two instances sharing release+environment+dist but differing in some - option are stored as one (first-seen) record, so genuine variation is silently lost. - -The optional hash exists to address exactly that last weakness. - -#### The optional SDK-set options hash - -SDKs **MAY** additionally set the `options_hash` field on the `sdk_config` payload — a **hash of -their options that is stable across instances sharing the same configuration** (same config → same -hash; the hash changes when the config changes). When a payload carries this hash, **the server -includes it in the dedup key**; when it is absent, dedup falls back to the composite key alone. - -Including the hash means two instances with the same release+environment+dist but genuinely -different options (different hash) are stored as **distinct records** — so within-release variation -and drift become visible instead of being collapsed into the first-seen config. - -- **The hash is computed off the serialized options block as defined - in the [serialization rules](#options-serialization-rules) — which makes it deterministic and - stable: it is **identical across instances that share a configuration** and **differs when the - configuration differs**. It does not need to be comparable _across_ SDKs, so the specific hash - algorithm is up to each SDK. -- **The hash is computed SDK-side, by design.** Otherwise, we cannot reliable relate events to their respective config. -- **If a hash is used, it MUST also be attached to every event the SDK produces**, so events can be - correlated back to the exact config that produced them: - - as an **attribute** on spans, logs, and other attribute-carrying items. We propose `sentry.config_hash` as a semantic attribute. - - as a **context field** on error and transaction events. We propose `sdk_config.hash` as a new context with a single field for now. + differences _within_ a tuple: instances sharing release+environment+dist but differing in some + option resolve to one (first-seen) record, so genuine variation is not distinguishable at this + level. This is exactly why the hash is the primary key and this is only the fallback. #### Correlating an event to its config This falls directly out of the dedup key: -- **Without a hash:** correlate via the composite key the event already carries +- **By hash (primary):** correlation is **exact** — the event's stamped hash matches exactly one + stored `sdk_config` record. This is the whole reason the hash must also live on events. +- **By composite key (fallback):** when no stored config matches the event's hash — or the event + carries no hash yet (see Timing below) — correlate via the composite key the event already carries (`release` + `environment` + `dist`). This is free and needs no event changes, but is only **bucket-level** — if config varied within the tuple, the event resolves to the bucket (a single first-seen record, or the set of stored variants), not necessarily the exact config that produced it. -- **With a hash:** correlation is **exact** — the event's stamped hash matches exactly one stored - `sdk_config` record. This is the whole reason the hash must also live on events. #### Trade-offs and caveats of the hash @@ -710,10 +717,10 @@ This falls directly out of the dedup key: likely should **not** be billed like events — but that needs to be decided explicitly. - **Outcomes / observability.** How are dropped or rejected `sdk_config` items recorded (outcomes, reasons) so we can see ingestion health for this new type? -- **Finalizing the dedup key.** (See [Storing](#storing) above.) We recommend - `release` + `environment` + `dist` plus an optional SDK-set options hash, but the exact field set, - the fallback for the common release-less case, and what SDKs should hash (and how they keep it - stable) still need to be nailed down. +- **Finalizing the dedup key.** (See [Storing](#storing) above.) We recommend a required SDK-set + options hash as the primary key, with `release` + `environment` + `dist` as the fallback lookup, + but the exact field set, the fallback behavior when no stored config matches an event's hash, and + what SDKs should hash (and how they keep it stable) still need to be nailed down. # Drawbacks @@ -728,10 +735,10 @@ This falls directly out of the dedup key: ingestion load, a new storage model, and server-side processing that did not exist before. - **Data may be incomplete or misleading.** The techniques that keep volume down also reduce fidelity: client sampling can under-sample or miss low-traffic releases and rare - configurations, and when no options hash is set, dedup keeps only the first-seen config per - `release`+`environment`+`dist` bucket even when config actually varies within it. Decisions made - on this data (e.g. deprecating an option that "looks unused") could be based on a - non-representative picture. + configurations, and when an event has to fall back to the composite key (no stored config matches + its hash), it resolves only to the first-seen config per `release`+`environment`+`dist` bucket even + when config actually varies within it. Decisions made on this data (e.g. deprecating an option that + "looks unused") could be based on a non-representative picture. - **Lossy representation of runtime options.** Reducing callbacks to `"[Function]"` tells us a `beforeSend`/`tracesSampler` exists but nothing about what it does. For audit-style use cases ("warn about confusing behavior") this is a hard limit — we can see that filtering is @@ -743,11 +750,10 @@ This falls directly out of the dedup key: - **Added SDK complexity and runtime cost.** Every SDK gains new machinery: debounced sending, flush-on-shutdown, opt-in per-integration `applied` tracking, normalization, and (for clients) sampling. This is more code, more surface for bugs, and some runtime overhead on every init. -- **Unresolved dedup identity.** Storage relies on a good key to deduplicate by. The composite key - (`release`+`environment`+`dist`) is coarse when `release` is absent, and the finer-grained - optional options hash is only as good as each SDK's hashing (and is not always present). Getting - this wrong means either storing too much or collapsing genuinely different configurations - together (see Storing). +- **Unresolved dedup identity.** Storage relies on a good key to deduplicate by. The required options + hash is only as good as each SDK's hashing, and the composite fallback key + (`release`+`environment`+`dist`) is coarse when `release` is absent. Getting this wrong means + either storing too much or collapsing genuinely different configurations together (see Storing). # Not in scope / Follow-up work From f2c9a46e53388b75f7c1297eab5c695f258e5737 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 12:52:45 +0200 Subject: [PATCH 27/51] add version --- text/0162-capture-sdk-options.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index a3c18352..82d22384 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -211,6 +211,7 @@ The shape below is the **stored** payload. SDKs send everything here **except** ```json { + "version": 1, "timestamp": "2026-09-21T12:00:00Z", "sdk": { @@ -275,6 +276,13 @@ The shape below is the **stored** payload. SDKs send everything here **except** ## Fields +- **`version`** — the schema version of the `sdk_config` payload itself, as an integer starting at + `1`. This is distinct from the SDK version (`sdk.version`) and the item-type name; it versions the + _shape_ described here. It lets us evolve the payload — add, rename, or restructure fields, or + change the serialization/normalization rules — in a way the server can detect and handle + explicitly (route to the right parser, migrate on read, or reject unknown majors) rather than + guessing from the field set. We bump it only for changes that a consumer must know about; + purely additive, forward-compatible fields (e.g. new `_other` keys) do not require a bump. - **`timestamp`** — when the payload was generated/sent by the SDK. We need to know when a configuration was reported, both to order records and to track configuration changes over time (e.g. "you changed your filtering rules in May"). Format follows the existing Sentry From 42a9c876d32a41d4f7db68c7bd5067f52d96fd42 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 13:00:14 +0200 Subject: [PATCH 28/51] adjustments to sending --- text/0162-capture-sdk-options.md | 36 ++++++++++++++++++++++++++++++++ 1 file changed, 36 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 82d22384..a4e3e627 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -569,6 +569,15 @@ reports, if applicable. If the process dies before the first flush opportunity, simply not sent for that invocation — acceptable given these instances are typically numerous and short, and an equivalent instance will report. +**Periodic re-send for long-lived processes.** Sending the config exactly once per init assumes the +one send reliably lands. That may not hold: if Relay **cannot guarantee that a `200` response means +the `sdk_config` was actually accepted and persisted** (e.g. it is dropped later in the pipeline, lost +to a transient downstream failure, sampled, or rate-limited after the ack), a long-lived server that +sent once at startup and then ran for days would show **no config at all** — a single lost send +becomes a permanent gap for the whole lifetime of that process. To bound that risk, server SDKs +**MAY periodically re-send** the (unchanged) config on a slow cadence — e.g. **once an hour** — so a +missed send self-heals on the next interval instead of being lost until the next restart. + ## Sending — Client SDKs (browser, mobile, gaming) Client SDKs are trickier. Unlike a server process — where one long-lived instance can send a @@ -603,6 +612,33 @@ small sample is enough to reconstruct the config for a release. configuration) may be under-sampled or missed entirely. It also adds configuration surface and is harder to reason about and test than a deterministic approach. +### Option III: Piggyback on the first outgoing envelope + +Rather than sending the `sdk_config` payload as its own request, **attach it as an additional item on +the first envelope the SDK sends anyway** (the first error, transaction, span, log, etc.). No traffic +means no `sdk_config`; as soon as the instance sends its first real payload, the config rides along in +the same envelope. This is a cross-cutting transport strategy — it can be combined with the +send-on-every-init or sampling decision above, and it applies equally to **server SDKs** (as an +alternative to the standalone debounced send). + +- **Benefits:** **Saves an extra request** — the config is delivered on an envelope that was already + going out, adding no network round-trip. It also naturally scopes reporting to instances that + actually produce data (an instance that never sends anything never sends its config either, which is + often exactly what we want). +- **Disadvantages:** + - **Trickier to implement.** The SDK has to hold the (possibly not-yet-finalized) config and hook + into envelope construction to append it to the first outgoing envelope, rather than just + scheduling a self-contained send. + - **The first event can fire before the config has settled.** If the first error/span/log is + emitted very early — before the debounce/settle window (see the `applied` and async-detection + timing above) — the SDK either has to attach an incomplete snapshot, or delay/skip attaching to + that first envelope and wait for a later one, which erodes the "one extra-free send" benefit. + - **Instances that never send an event are invisible.** If nothing is ever sent — including the + case where a **misconfiguration prevents any event from being sent at all** — the config is never + reported. That is precisely one of the situations we most want visibility into ("is this set up + wrong?"), and this approach cannot surface it. A standalone send (Option I/II) does not have this + blind spot. + ### Serverless Although serverless functions run "server" SDKs, they share the key constraints of client From a97297db8a118541c0bf2179e45ca4b82ccdc324 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 13:03:50 +0200 Subject: [PATCH 29/51] add eap note --- text/0162-capture-sdk-options.md | 8 ++++++++ 1 file changed, 8 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index a4e3e627..652fc455 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -274,6 +274,14 @@ The shape below is the **stored** payload. SDKs send everything here **except** } ``` +> **Note on EAP storage.** EAP (our events analytics platform) prefers a single top-level +> `attributes` field with everything nested inside it, rather than the multiple structured top-level +> blocks shown above. If we need to conform to that, we would have to adapt this shape into a semantic +> attribute shape with JSON-valued fields — e.g. `options` becomes one attribute (a JSON blob), +> `integrations` another, and so on — and make sure we can still **query effectively** across those +> JSON attributes (filtering on a specific option, a normalized key, or the options hash). Whether we +> store in EAP and thus adopt this shape is an open question to resolve with the ingest/storage design. + ## Fields - **`version`** — the schema version of the `sdk_config` payload itself, as an integer starting at From b9e5a454b3c5571d888238a5882e93150aa41e04 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 13:07:18 +0200 Subject: [PATCH 30/51] more name ideas --- text/0162-capture-sdk-options.md | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 652fc455..d9f1a6c6 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -135,9 +135,10 @@ snake_case convention for item types (`event`, `transaction`, `client_report`, ` `sdk_config` reads well because the payload is broader than just the raw `init()` options — it also carries SDK identity, integration status, and general metadata. -Alternatives considered: `sdk_options` (closest to the literal `init()` arguments, but narrower -than what the payload actually contains) and `client_config` (risks confusion with Sentry -"client reports" and with the SDK's internal `Client`). We recommend `sdk_config`. +Ideas for envelope item name, TBD: +* `sdk_config` +* `sdk_metadata` +* `sdk_options` **Backward compatibility with older ingest / self-hosted.** Introducing a new envelope item type is safe for older infrastructure. Envelopes are designed so that unknown item types are simply From ec2c5cec7654437670dc59931c3222459e3d9c7b Mon Sep 17 00:00:00 2001 From: Daniel Szoke <7881302+szokeasaurusrex@users.noreply.github.com> Date: Fri, 25 Sep 2026 14:31:40 +0200 Subject: [PATCH 31/51] ref: Shorten RFC 0126 significantly (#163) --- text/0162-capture-sdk-options.md | 956 +++++++------------------------ 1 file changed, 221 insertions(+), 735 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index d9f1a6c6..750cad0f 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -8,207 +8,66 @@ # Summary Today, we have very little visibility into which options a given Sentry SDK instance was configured -with. This RFC proposes a mechanism for SDKs to report the configuration they were -initialized with (the arguments passed to `Sentry.init()`, plus relevant derived/effective -values) to Sentry, so that this information can be stored, surfaced, and acted upon. - -Knowing the configured options unlocks a range of use cases — from self-healing and -debugging ("was this simply not enabled, or did the event never get sent?"), to product -analytics (which options are actually used, which matter most), to user-facing features -(showing all `Sentry.init()`s producing data into a project, auditing setups, and eventually -allowing configuration changes from the UI). We do not need to build all of these up front; -this RFC focuses on the capture and transport mechanism, with the downstream use cases -described as motivation and future work. +with. This RFC proposes that SDKs report their configuration (the options passed to `Sentry.init()`, +plus relevant derived and effective values) in a new `sdk_config` envelope item, so that Sentry can +store, surface, and act on it. Product features built on this data are future work. # Motivation -We currently cannot answer basic questions about how an SDK instance is configured. This -creates gaps in several areas: - -- **Self-healing / support**: When data is missing or filtered, we cannot tell whether a - feature was disabled, a sample rate dropped the event, or something failed. Knowing the - configured options would let us distinguish "never enabled" from "enabled but filtered". -- **Analytics & product decisions**: We do not know which options are actually used in the - wild. This data would inform what to highlight in docs, what to deprecate or remove in - major versions, and where to invest. -- **Configuration-aware querying**: Configuration changes can affect trends over time (e.g. - a change to filtering or sampling rules). Surfacing these changes would help explain shifts - in data. -- **Setup audits & warnings**: We could audit `init()` setups to suggest improvements or warn - users about potentially confusing behavior given their settings. -- **Discoverability of data sources**: Users could see all of the `Sentry.init()`s producing - data into a project, understand the sources of their data, and see the filtering, sampling, - and other configuration that affects it — including changes over time. +A user reports missing transactions. Was tracing never enabled? Did a sample rate or +`beforeSendTransaction` drop them? Did something fail? Without the SDK configuration, we cannot tell. +With it, we can support: + +- **Self-healing and support:** distinguish "never enabled", "enabled but filtered", and "failed". +- **Analytics and product decisions:** learn which options are used and which matter most, to guide + docs, deprecations and removals in major versions, and investment. +- **Configuration-aware querying:** explain shifts in data with configuration changes, such as new + filtering or sampling rules. +- **Setup audits and warnings:** suggest improvements and warn about confusing behavior given the + settings. +- **Discoverability of data sources:** show every `Sentry.init()` that sends data into a project, with + its configuration over time, and eventually allow changing configuration from the UI. # Background -Today we do not capture the configuration of an SDK instance in any meaningful, first-class -way. There is no place in Sentry where you can look up "what options was this -`Sentry.init()` called with", and consequently none of the use cases in the Motivation are -possible today. - -There are, however, a few adjacent things that already exist. They each capture a small -slice of related information, but none of them gives us the configured options, and none is -designed for that purpose: - -- **SDK metadata on error/transaction events.** Events carry an `sdk` object (name, version, - and lists of `integrations` and `packages`), and some settings leak into events - indirectly. This tells us _which integrations are present_ and the SDK version, but not - _how_ things were configured (sample rates, `beforeSend`, transport options, `debug`, - `environment` defaults, `sendDefaultPii`, denyUrls/allowUrls, tracing options, and so on). - It is also only present when an event is actually sent — so an instance that is configured - but never produces an event (or whose events are all filtered) is invisible. This is the - closest existing signal, but it is partial and event-coupled. - -- **Client reports.** Client reports are sent as their own, separate envelope item - (`client_report`), independent of any error or transaction event. Their _content_ is - unrelated to configuration — they report aggregate counts of discarded events by `reason` - and `category` (e.g. rate-limited, sample-rate, before-send). But they are a useful - precedent for _how_ we might send configuration: they show that we already have a pattern - for the SDK to emit a standalone, non-event payload on its own cadence (batched/periodic, - flushed on shutdown). A "SDK configuration" report could plausibly follow a similar - transport shape rather than being attached to individual events. - -In short: what we have today is either a partial, event-coupled snapshot (SDK metadata) or a -transport precedent with unrelated content (client reports). Neither captures the configured -options, which is what this RFC is about. +Sentry has no first-class record of SDK configuration. Two existing mechanisms come close: + +- **The `sdk` object on error and transaction events** lists the SDK name, version, integrations, and + packages, but not how the SDK is configured (sample rates, `beforeSend`, `sendDefaultPii`, + `denyUrls`, transport and tracing options, …). It only arrives with events, so instances that send + none, or whose events are all filtered, are invisible. +- **Client reports** count discarded events, which is unrelated to configuration. They are, however, + a precedent for a standalone, non-event item that SDKs send on their own cadence and flush on + shutdown. # Options Considered -Broadly, there are two ways to get configuration data from the SDK to Sentry. Both assume -the SDK can produce a serialized view of its options; they differ in _how that view is -transported and handled_. - -## Option A (preferred): A dedicated envelope item for SDK options - -Introduce a new, first-class envelope item type dedicated to SDK configuration. The SDK -serializes its options and sends them as a standalone payload, handled on its own path -server-side rather than being coupled to error/transaction events. - -This is the preferred option because it decouples "what is this instance configured with" -from "did this instance send an event". It gives us a clean, purpose-built payload we can -version, store, and reason about independently, and it is not dependent on an error ever -being produced. It also follows the transport precedent set by client reports (a separate, -non-event envelope item on its own cadence; see Background). - -With this option, the design work is mostly about two questions, which we will dive into in -more detail: - -- **a) The shape of the envelope** — what exactly we put into the payload: which options we - capture, how we represent non-serializable options (functions like `beforeSend`, - integration instances), how we normalize/redact sensitive values, and how we keep the - schema consistent and generalizable across SDKs (this starts with JS). -- **b) How/when to send it, and how/when to store it** — the send cadence (once per init, on - change, periodically, flushed on shutdown), and the server-side handling: where it lands, - how we group instances, how we deduplicate identical configs, and how we track changes over - time. - -## Option B: Expand error events to carry SDK options - -Alternatively, we could piggyback on error (and transaction) events — for example by -expanding the existing `sdk` key on the event to carry the full configured options, and then -doing the work server-side to infer, group, and store this out of the event stream. - -This avoids a new envelope type and reuses an existing, well-understood transport. However, -it inherits the drawbacks of being event-coupled: configuration is only observed when (and -as often as) events are sent, an instance that never produces an event is invisible, -identical config is re-sent on every event (wasteful, needs server-side dedup), and we -overload the event schema and pipeline with data that is not really about the event. The -server-side grouping/storage problem also becomes harder because the signal is buried inside -the high-volume event stream. - -For these reasons Option A is preferred; the remaining sections focus on the questions it -raises. - -# Envelope Shape - -This section proposes the shape of the dedicated SDK-options payload (question **a** above). -The goal is a **language-agnostic** shape that works equally for JavaScript, Python, and every other SDK, so the server can handle a single, consistent schema. - -## Envelope item type - -We propose naming the new envelope item type **`sdk_config`**, following the existing -snake_case convention for item types (`event`, `transaction`, `client_report`, `session`, …). -`sdk_config` reads well because the payload is broader than just the raw `init()` options — it -also carries SDK identity, integration status, and general metadata. - -Ideas for envelope item name, TBD: -* `sdk_config` -* `sdk_metadata` -* `sdk_options` - -**Backward compatibility with older ingest / self-hosted.** Introducing a new envelope item type -is safe for older infrastructure. Envelopes are designed so that unknown item types are simply -**ignored**: an older Relay / self-hosted Sentry that predates `sdk_config` will not recognize the -item and will **discard it**, while still processing the other items in the same envelope (errors, -transactions, etc.) as usual. That is the desired behavior here — SDKs can start emitting -`sdk_config` unconditionally, and setups that cannot yet handle it lose only this -new-and-supplementary payload, with no impact on existing data. No SDK-side version gating is -required. - -## Design principles - -- **Primitives only.** The payload contains only JSON-serializable values: strings, numbers, - booleans, `null`, arrays, and plain objects. -- **Callbacks and other runtime values are reduced to markers.** We do not care about a - callback's implementation, only that _a user-defined callback was set_. Any non-serializable - value is normalized to a sentinel, following the exact rules in - [Options serialization rules](#options-serialization-rules) below (functions → `"[Function]"`, - other runtime constructs → a type marker). -- **Effective config, natively named but flat.** The `options` block carries the SDK's final, - effective options (after defaults and derivation), keyed by native option names. Nested objects - are **flattened with dot notation** (e.g. `dataCollection.http.bodies`) rather than kept as nested - objects, so the structure stays flat and uniform (see the serialization rules). Which of those - keys the user explicitly set is recorded separately in `options_set_by_user`. -- **Generic representation of integration options.** Integrations are user-configurable - (`Sentry.myIntegration({ filter: 'aaa' })`), so we must capture their options generically — - keyed by integration name, with their options normalized by the same rules. We do not need - to understand any specific integration's options; we just record them. -- **Room for well-known metadata and for open-ended data.** Alongside the raw options there is - space for well-known, structured metadata (SDK identity, release/environment, etc.) and a - free-form bucket SDKs can use for anything not yet modeled. - -## Options serialization rules - -When serializing the `options` block (and each integration's `options`), SDKs apply the following -rules, top to bottom, to every value. The goal is a deterministic, primitives-only representation -that is consistent across SDKs. - -- **Primitives pass through.** Strings, numbers, booleans, and `null` are emitted as-is. Arrays are - emitted as arrays, with each element serialized by these same rules. -- **Nested objects are flattened with dot notation.** We deliberately keep the structure **flat**: - a nested option object is not emitted as a nested object but as dotted keys joining the path, e.g. - `dataCollection: { http: { bodies: true } }` becomes `"dataCollection.http.bodies": true`. The - same applies inside each integration's `options`. This gives a single, flat, uniform key space - that is easy to store, query, and normalize; the leaf values are serialized by the other rules - here. (Arrays are not flattened — they are kept as arrays.) -- **Functions become `"[Function]"`.** Any callback (`beforeSend`, `tracesSampler`, - `beforeBreadcrumb`, transport factories, etc.) is replaced by the literal string `"[Function]"`. - SDKs MAY include the function name when readily available (`"[Function: beforeSend]"`), but the - bare `"[Function]"` marker is the required baseline — consumers must not depend on the name. -- **Integrations are omitted from `options` entirely.** The `integrations` option is not emitted in - the `options` block at all — everything we capture about integrations (their identity, runtime - status, and own configured options) lives in the dedicated top-level `integrations` block, keyed by - integration name. So `Sentry.init({ integrations: [Sentry.myIntegration({ filter: 'aaa' })] })` - yields no `integrations` key in `options`, and in the `integrations` block - `"MyIntegration": { "options": { "filter": "aaa" } }`. This avoids duplicating the integration - data in two places. -- **Other runtime constructs become a type marker.** Any remaining non-serializable value (a class - instance, stream, socket, etc.) is replaced by a bracketed type marker, e.g. `"[SomeType]"`, - reusing each SDK's existing normalization convention (e.g. JS `normalize()`). - -These rules are what "normalized" means throughout this document. Scrubbing sensitive values is a -**separate, primarily server-side concern** (see `options`) — SDKs **MAY** scrub values they know -to be sensitive, but they do not attempt to guarantee fully-scrubbed data. - -## Proposed shape - -The shape below is the **stored** payload. SDKs send everything here **except** -`normalized_options`, which Relay derives at ingestion time from `options` (see -[Normalization at ingestion](#normalization-at-ingestion)). +## Option A (preferred): A dedicated envelope item + +SDKs send the configuration in a standalone `sdk_config` envelope item, like client reports. This +decouples "how is this instance configured?" from "did it send an event?", and gives us a payload +that we can version, store, and analyze independently. + +## Option B: Add the options to error events + +Extend the `sdk` object on error and transaction events with the full options. This reuses the +existing transport, but configuration is only seen when events are sent, identical configuration is +re-sent with every event, and the data is mixed into the high-volume event stream, which makes +grouping and storing it harder. + +# Proposed Design + +## Payload + +The item type is tentatively named `sdk_config` (alternatives: `sdk_metadata`, `sdk_options`), with +one language-agnostic schema for all SDKs, starting with JavaScript. Older Relay and self-hosted +versions discard unknown item types without affecting the rest of the envelope, so SDKs can send +`sdk_config` unconditionally. + +SDKs send every field in this example except `normalized_options`, which Relay adds at ingestion: ```json { @@ -275,552 +134,179 @@ The shape below is the **stored** payload. SDKs send everything here **except** } ``` -> **Note on EAP storage.** EAP (our events analytics platform) prefers a single top-level -> `attributes` field with everything nested inside it, rather than the multiple structured top-level -> blocks shown above. If we need to conform to that, we would have to adapt this shape into a semantic -> attribute shape with JSON-valued fields — e.g. `options` becomes one attribute (a JSON blob), -> `integrations` another, and so on — and make sure we can still **query effectively** across those -> JSON attributes (filtering on a specific option, a normalized key, or the options hash). Whether we -> store in EAP and thus adopt this shape is an open question to resolve with the ingest/storage design. - -## Fields - -- **`version`** — the schema version of the `sdk_config` payload itself, as an integer starting at - `1`. This is distinct from the SDK version (`sdk.version`) and the item-type name; it versions the - _shape_ described here. It lets us evolve the payload — add, rename, or restructure fields, or - change the serialization/normalization rules — in a way the server can detect and handle - explicitly (route to the right parser, migrate on read, or reject unknown majors) rather than - guessing from the field set. We bump it only for changes that a consumer must know about; - purely additive, forward-compatible fields (e.g. new `_other` keys) do not require a bump. -- **`timestamp`** — when the payload was generated/sent by the SDK. We need to know when a - configuration was reported, both to order records and to track configuration changes over - time (e.g. "you changed your filtering rules in May"). Format follows the existing Sentry - convention (ISO 8601 shown here; could equally be epoch seconds to match event `timestamp`). -- **`sdk`** — SDK identity metadata: `name`, `version`, and `packages`. This is the same - information SDKs attach to error/transaction events today; **this RFC proposes moving it - here** so it lives in one canonical place. (The set of integrations, previously part of the - event's `sdk` object as a list of names, is captured more richly in the top-level - `integrations` block below.) -- **`integrations`** — a map keyed by integration name, where each entry holds everything we know - about that integration in one place: its runtime status and its configured options. This - consolidates what would otherwise be two parallel name-keyed maps. Per entry: - - The **presence of the key** means the integration is registered/enabled. - - An optional, opt-in **`applied`** boolean records whether the integration determined at runtime - that it actually took effect (see "Reflecting which integrations are actually used"). - - **`options`** is the integration's own normalized options, nested under this key so arbitrary - user options can never collide with well-known status keys like `applied`. Integrations with no - options report `"options": {}`. This is how we generically capture things like - `MyIntegration({ filter: 'aaa' })` without understanding any specific integration. -- **`meta`** — well-known, general metadata that we want first-class regardless of how it was - set: `release`, `environment`, `dist`, and runtime/platform information (e.g. runtime name - and version). These describe the instance producing data, complementing the raw `options`. -- **`options`** — a normalized snapshot of the SDK's **final, effective** configuration (i.e. the - options object the SDK actually runs with, after defaults, env-var resolution, and any - derived/integration-injected values), keyed by native names and flattened with dot notation, with - all values reduced to primitives per the rules above. Sending the effective config — rather than only the literal - `init()` arguments — is what lets us answer behavior questions ("what sample rate is actually in - effect", "is it enabled"); it also reads directly off the SDK's existing options object with no - extra plumbing, and reflects the settled state at send time (see the debounce in Sending). `integrations` are omitted here. - Sensitive data - (especially tokens and other secrets) is **primarily scrubbed server-side**; SDKs **MAY** - additionally scrub values they know to be sensitive (e.g. a field that always holds a secret), - but this is a best-effort defense-in-depth measure — SDKs do **not** attempt to guarantee - fully-scrubbed data, and the server-side scrubbing remains the mechanism we rely on. -- **`options_set_by_user`** — a flat array of the **native option keys the user explicitly set** - in `init()` (as opposed to values that came from defaults, env vars, or integrations). This is - the signal that lets us distinguish default values from user-set values - — essential for adoption analytics and setup audits, where a default value is not - "usage". This is effectively similar to `Object.keys(options)` where `options` are the user-provided options for `Sentry.init(options)`. - Keys here use the **same flattened dot-notation as `options`**, so a user-set nested value is - listed by its dotted leaf key (e.g. `dataCollection.http.bodies`) and always corresponds 1:1 to a - key present in `options`. -- **`options_hash`** — an SDK-computed hash of the options that is **stable across - instances sharing the same configuration** (same config → same hash). SDKs **MUST** compute and - send it, and **MUST** stamp the **same value on every event** (attribute on spans/logs, context - field on errors/transactions) so events correlate exactly to their config. It is the fine-grained - component of the dedup key: the server folds it in so genuinely different configs within the same - `release`+`environment`+`dist` are stored as distinct records. See [Storing](#storing) for the - full rules and caveats, including the correlation fallback when no matching stored config exists. - It is computed off the serialized `options` & `integrations` blocks (both normalized - by the same [serialization rules](#options-serialization-rules)), so a change to either — including - an integration's own options or its `applied` status — produces a different hash. This makes it - deterministic and stable across instances; it does not need to be comparable across SDKs. -- **`normalized_options`** — a **Relay-derived** subset of `options`, keyed by canonical - cross-SDK names, produced at ingestion (see below). SDKs never send this block. Each entry maps - a canonical key to `{ "key": , "value": }`, so consumers - can compare the same option across SDKs while still seeing what it was called natively. Whether - the user set it is answered the same way as for any other option — via `options_set_by_user`. -- **`_other`** — a free-form, SDK-defined bucket for anything not covered by the well-known - fields above. The leading underscore signals that this is arbitrary, unstructured data. - Keeps the schema forward-compatible: SDKs can record additional data without a schema change, - and useful keys can later be promoted to first-class fields. - -## Why effective options plus a flat set-by-user array - -We deliberately split configuration into two SDK-sent pieces — the full **effective `options`** -and a flat **`options_set_by_user`** array — rather than, say, shipping both a user-provided and an -effective options tree, or wrapping every option value in a `{ value, source }` object. The -benefits: - -- **Easy to implement in SDKs.** `options` is essentially the SDK's existing effective options - object, serialized — most SDKs can read it straight from an existing API (e.g. JS - `client.getOptions()`, or the equivalent effective-options accessor in other SDKs), so there is - no second config snapshot to capture or keep around. `options_set_by_user` is produced by a - single diff of the raw `init()` argument's keys against that object. No per-option plumbing, no - wrapper types, no bookkeeping threaded through the option system. -- **Easy to reason about.** `options` means exactly one thing — the configuration the SDK actually - runs with — and `options_set_by_user` means exactly one thing — which of those the user chose. - There is no ambiguity about whether a given block is "before" or "after" defaults, and no mixed - value/metadata shape to interpret. What each field represents is obvious from its name. -- **Can be joined as needed.** Keeping provenance as a separate flat set means any consumer can - answer "was this user-set?" for _any_ option with a simple membership check, and can just as - easily ignore provenance entirely when it does not care. The two pieces compose on demand - (including for `normalized_options`, which joins against the same array) instead of being - pre-fused into one heavier structure that every consumer pays for whether or not they need it. - -The trade-off is that provenance is not co-located with each value (you look it up rather than -reading it inline), and the array is top-level-only. Both are acceptable given how much simpler -this keeps the SDK side and the schema. - -### Known limitation: user-provided _form_ is not preserved - -Because we send the **effective** options, we lose information when the user expressed a value in a -form that the SDK coerces into something else before it lands on the effective options. The value -we report is the coerced result, and `options_set_by_user` only tells us the option _was_ set — not -_how_ it was written. Concretely: a user can pass `integrations` as either an array or a -**function** (`(defaults) => Integration[]`), but the client always ends up holding a resolved -array — which is what the top-level `integrations` block reflects — so we cannot tell, from the -payload, whether the user configured integrations via a function. We therefore cannot answer -questions like "how many users pass `integrations` as a function". - -We accept this limitation for now: - -- **It is rare.** In the JS SDK, the option type system pins genuine user-vs-effective _shape_ - divergence to essentially two options: `integrations` (array-or-function → array) and - `stackParser` (array-or-function → function). Everything else keeps the same shape; only defaults - or env values get filled in, which the effective `options` already captures faithfully. -- **`integrations` is the main case that matters**, and even there the question ("was it a - function?") is a nice-to-have, not core to the primary use cases. - -If we later decide this signal is worth capturing, we can layer it on **without reworking the -shape** — e.g. an optional, sparse `options_user_provided` block that records the user-provided -(serialized) value _only_ for the few keys whose provided form differs from the effective one (so -it would carry `{ "integrations": "[Function]" }` and otherwise be empty). We deliberately leave -that out of the initial design and revisit it only if a concrete need arises. - -## Reflecting which integrations are actually used - -Knowing which integrations are _registered_ is not the same as knowing which are actually -_doing anything_. In the Node SDK, for example, a large set of integrations is added by -default (`ExpressIntegration`, `FastifyIntegration`, `KoaIntegration`, …), but a given app -typically uses only one of them. For analytics and audits we care about the difference -between "this integration is present because it ships by default" and "this integration is -actually instrumenting this app". - -This is captured generically, without enumerating any specific integration, via the `applied` key -on each entry of the top-level `integrations` map: - -- The **presence of a key** means the integration is registered/enabled. -- An optional **`applied`** boolean means the integration determined at runtime whether it - actually took effect. `ExpressIntegration` sets `applied: true` once it successfully patches - Express; a defaulted integration whose target framework is absent can report - `applied: false`. An integration that reports nothing simply omits `applied` (absent / unknown). - -Key properties of this design: - -- **Generic.** No integration-specific fields in the schema; any integration can contribute - the well-known `applied` signal (and we can add further opt-in status keys later). -- **Opt-in and non-exhaustive.** Integrations are not required to report status. We selectively - push this into the integrations where the signal is valuable to us (e.g. the framework - integrations), and leave the rest unreported. -- **Namespaced from options.** `applied` (and any future status key) sits at the top of the entry, - while the integration's own configured options are nested under the entry's `options` key, so - the two never collide. - -Note there is a timing implication: `applied` is often only known slightly after `init()` (once -instrumentation runs), which influences _when_ the payload is sent or updated. This is -discussed in the send/store section. - -## Cross-SDK naming - -Option keys differ across SDKs (JS `tracesSampleRate` vs. Python `traces_sample_rate`). To -compare the same option across languages we want a canonical cross-SDK vocabulary (in the spirit -of [0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md)), but we do **not** -want every SDK to own that mapping and keep it in sync. - -**Recommendation: SDKs send native names; Relay normalizes a defined subset.** The `options` -block keeps each SDK's **native option names as-is** — SDKs report options exactly as they are -named in that SDK, with no normalization to a shared vocabulary. This keeps the SDK side dumb and -maintenance-free: there is nothing to map, nothing to keep in sync, and new options are captured -automatically. The canonical mapping lives **server-side in Relay**, which reads a fixed set of -well-known options out of `options` and emits them into `normalized_options` under canonical keys -(see below). - -This gets us the best of both: full, native fidelity in `options`, plus a normalized, -cross-SDK-comparable view in `normalized_options` — without pushing catalog upkeep into every SDK. - -**Why Relay, not the SDK.** Doing the normalization server-side is a deliberate choice, for -several reinforcing reasons: - -- **Simpler SDKs.** Each SDK only has to serialize and send its own native options (something it - effectively already has). It does not need to know the canonical vocabulary, map its keys onto - it, or reason about equivalence across languages — that logic never ships in the SDK at all. -- **Nothing to keep aligned across SDKs.** A canonical mapping owned by the SDKs would have to be - implemented, and kept consistent, in _every_ SDK and language independently. Any drift (a - mismatched canonical key, a missed option, an inconsistent value normalization) would silently - corrupt cross-SDK comparisons. Centralizing removes that entire class of cross-SDK - synchronization problem. -- **One place to implement and maintain.** The normalization logic and the knowledge of which - canonical keys exist live in a **single system (Relay)** rather than being duplicated across N - SDKs. There is one implementation to write, test, review, and reason about — not one per SDK. -- **Changeable over time without SDK releases.** What we choose to normalize will evolve. Because - the catalog lives in Relay, we can add, rename, or refine normalized keys — and re-normalize - already-ingested `options` — as a Relay change alone, with **no SDK update, release, or user - upgrade** required. If normalization lived in the SDKs, every change would mean shipping every - SDK and waiting for the ecosystem to upgrade, so the normalized dataset would always lag. - -## Normalization at ingestion - -`normalized_options` is produced by **Relay at ingestion time**, never sent by the SDK. For each -option in the normalization catalog, Relay looks it up in the incoming `options` (by the native -key registered for that SDK) and, if present, emits an entry: - -```json -"normalized_options": { - "traces_sample_rate": { "key": "tracesSampleRate", "value": 0.5 } -} -``` - -- The **outer key** is the canonical, cross-SDK name (snake_case). -- **`key`** is the native option name it was normalized from, so the original naming is not lost. -- **`value`** is the value of this option (the effective value from `options`). - -Provenance is intentionally not repeated here: to check whether a normalized option was user-set, -look up its native `key` in `options_set_by_user`, exactly as you would for any raw option. - -Only options in the catalog are normalized; everything else remains available under `options`. -Because the mapping is Relay-side (see [Why Relay, not the SDK](#cross-sdk-naming)), the catalog -can be extended without an SDK release, and already-stored raw `options` can be re-normalized. - -### Which options we normalize - -We start with a small, curated set of high-value options that are meaningful across SDKs. The -canonical key is the same across languages; only the native `key` differs. - -| Canonical key | Meaning | Native examples (JS → Python) | -| ---------------------- | -------------------------------- | ----------------------------------------- | -| `sample_rate` | Error sample rate | `sampleRate` → `sample_rate` | -| `traces_sample_rate` | Tracing sample rate | `tracesSampleRate` → `traces_sample_rate` | -| `profiles_sample_rate` | Profiling sample rate | `profilesSampleRate` → `profiles_sample_rate` | -| `send_default_pii` | Whether default PII is sent | `sendDefaultPii` → `send_default_pii` | -| `debug` | Debug logging enabled | `debug` → `debug` | -| `before_send` | Whether a `before_send` hook is set (marker) | `beforeSend` → `before_send` | -| `before_send_transaction` | Whether a `before_send_transaction` hook is set (marker) | `beforeSendTransaction` → `before_send_transaction` | -| `before_send_span` | Whether a `before_send_span` hook is set (marker) | `beforeSendSpan` → `before_send_span` | -| `before_send_log` | Whether a `before_send_log` hook is set (marker) | `beforeSendLog` → `before_send_log` | -| `ignore_spans` | Span-ignore rules | `ignoreSpans` → `ignore_spans` | -| `traces_sampler` | Whether a `traces_sampler` hook is set (marker) | `tracesSampler` → `traces_sampler` | -| `data_collection.*` | Data-collection settings (all flattened sub-keys) | `dataCollection.*` → `data_collection.*` | - -A trailing `.*` (e.g. `data_collection.*`) denotes a **family of flattened sub-keys**: every dotted -key under that path (`dataCollection.http.bodies`, `dataCollection.http.headers`, …) is normalized, -each becoming its own `normalized_options` entry under the canonical dotted key. - -`release`, `environment`, and `dist` are deliberately **not** in this catalog — they are already -promoted to first-class fields under `meta`. The catalog is intended to grow over time; the list -above is the initial, deliberately-conservative set, and the exact registry (canonical key ↔ -per-SDK native key) is maintained alongside Relay. - -# Sending and Storing - -This section covers question **b**: when the SDK emits the options payload, and how it is -handled server-side. When to send is **not one-size-fits-all** — it differs meaningfully -between SDK types. Broadly we distinguish **server SDKs** (Node, Python, Java, Go, …) from -**client SDKs** (browser, mobile, desktop, gaming, …), which have very different lifecycles. -We start with the server case; the client case is trickier and is covered next. - -## Sending — Server SDKs - -For server SDKs, the payload is sent **once, shortly after `init()`**, but only once the -configuration has **settled**. The hard requirement is only this: - -- **Send the final, settled options — not a half-initialized snapshot.** Some values are not - known at the exact moment `init()` returns — for example an integration's runtime `applied` - status (see "Reflecting which integrations are actually used"), lazily-registered integrations, - or release/environment detected asynchronously. The SDK must wait until it can reasonably expect - these to have settled before capturing and sending the payload. -- **One payload per init.** The goal is a single, stable payload per SDK instance/init, not a - stream of updates. If the config meaningfully changes later, that is handled as a separate - concern (and mostly matters for long-lived processes); the common case is one send per - process start. - -**How to wait is up to each SDK.** We deliberately do not mandate a single mechanism — each SDK -**MAY** use whatever hook or mechanism fits its runtime and lifecycle, as long as it reasonably -captures the final options. A few examples: - -- **A short debounce delay** (a reasonable default, e.g. on the order of **2 seconds**): schedule - the send a short time after `init()` rather than synchronously, letting the settling period - collapse into one payload. The delay **MAY be configurable** so setups that settle slower (e.g. - an SDK that only detects framework instrumentation after the first request) can tune it. -- **A lifecycle hook** the SDK already has — e.g. sending after the first request/transaction is - processed, on an "SDK ready"/post-init hook, or when the event loop first goes idle. -- **Any equivalent trigger** that reliably fires after the config the SDK cares about is known. - -Debouncing is the simplest baseline and a fine default, but it is an example, not a rule: an SDK -with a natural hook that guarantees settled options should prefer that. What matters is the -outcome — the payload reflects the effective configuration. - -**Short-lived processes.** Some server environments (serverless functions, short CLI -invocations) may exit before the debounce timer fires. In those cases the SDK should flush the -pending options payload on shutdown / at the same points it already flushes events, so the -config is not lost. SDKs MAY use the same or a similar approach as to how they flush client -reports, if applicable. If the process dies before the first flush opportunity, the payload is -simply not sent for that invocation — acceptable given these instances are typically -numerous and short, and an equivalent instance will report. - -**Periodic re-send for long-lived processes.** Sending the config exactly once per init assumes the -one send reliably lands. That may not hold: if Relay **cannot guarantee that a `200` response means -the `sdk_config` was actually accepted and persisted** (e.g. it is dropped later in the pipeline, lost -to a transient downstream failure, sampled, or rate-limited after the ack), a long-lived server that -sent once at startup and then ran for days would show **no config at all** — a single lost send -becomes a permanent gap for the whole lifetime of that process. To bound that risk, server SDKs -**MAY periodically re-send** the (unchanged) config on a slow cadence — e.g. **once an hour** — so a -missed send self-heals on the next interval instead of being lost until the next restart. - -## Sending — Client SDKs (browser, mobile, gaming) - -Client SDKs are trickier. Unlike a server process — where one long-lived instance can send a -single payload that represents a whole fleet's worth of traffic — client instances are -**numerous and short-lived**: every page load, app launch, or game session is its own `init()`. -Sending the options payload naively from each one would be dramatically more data than the -server case. Because of that, the _when_ needs more thought. We propose exploring the following -options. - -### Option I: Send on every init (with debounce) - -Do exactly what server SDKs do: send once per `init()`, guarded by the same short debounce. - -- **Benefits:** Easy and simple to implement and reason about. Identical mental model and code - path to the server case; no sampling logic, no extra configuration. Every instance is - represented. -- **Disadvantages:** Much higher overhead. Sentry has to ingest _much_ more data (one payload - per page load / app launch / session, at client-traffic volumes), and it adds a request for - users on every init. - -### Option II: Sampling - -Send the payload only for a random fraction of inits (e.g. **1%**, potentially configurable). -The intuition is that configuration is essentially **constant per release** — every instance of -the same release reports roughly the same options — so we do not need it from every init; a -small sample is enough to reconstruct the config for a release. - -- **Benefits:** Much less overhead on both sides, while still capturing the configuration for - each release given enough traffic. -- **Disadvantages:** Sampling is random, so it works better or worse depending on traffic - volume — a high-traffic release is well covered, a low-traffic one (or a rarely-hit - configuration) may be under-sampled or missed entirely. It also adds configuration surface and - is harder to reason about and test than a deterministic approach. - -### Option III: Piggyback on the first outgoing envelope - -Rather than sending the `sdk_config` payload as its own request, **attach it as an additional item on -the first envelope the SDK sends anyway** (the first error, transaction, span, log, etc.). No traffic -means no `sdk_config`; as soon as the instance sends its first real payload, the config rides along in -the same envelope. This is a cross-cutting transport strategy — it can be combined with the -send-on-every-init or sampling decision above, and it applies equally to **server SDKs** (as an -alternative to the standalone debounced send). - -- **Benefits:** **Saves an extra request** — the config is delivered on an envelope that was already - going out, adding no network round-trip. It also naturally scopes reporting to instances that - actually produce data (an instance that never sends anything never sends its config either, which is - often exactly what we want). -- **Disadvantages:** - - **Trickier to implement.** The SDK has to hold the (possibly not-yet-finalized) config and hook - into envelope construction to append it to the first outgoing envelope, rather than just - scheduling a self-contained send. - - **The first event can fire before the config has settled.** If the first error/span/log is - emitted very early — before the debounce/settle window (see the `applied` and async-detection - timing above) — the SDK either has to attach an incomplete snapshot, or delay/skip attaching to - that first envelope and wait for a later one, which erodes the "one extra-free send" benefit. - - **Instances that never send an event are invisible.** If nothing is ever sent — including the - case where a **misconfiguration prevents any event from being sent at all** — the config is never - reported. That is precisely one of the situations we most want visibility into ("is this set up - wrong?"), and this approach cannot surface it. A standalone send (Option I/II) does not have this - blind spot. - -### Serverless - -Although serverless functions run "server" SDKs, they share the key constraints of client -SDKs: instances are **numerous and short-lived**, with an `init()` per invocation rather than -one long-lived process. We therefore expect similar trade-offs to apply, and the same -solutions explored here for client SDKs (e.g. sampling) **COULD** be applied to serverless -environments as well, rather than the plain debounced send-on-every-init used for long-lived -server processes. - -## Storing - -Once payloads arrive, we have to decide how much of this data we actually keep. Broadly there -are two options. - -### Option 1: Store every record - -Persist every payload we receive as its own record. - -- **Benefits:** Complete history; nothing is lost, and we could in principle see every - individual instance's config over time. -- **Disadvantages:** Very high storage cost and volume, especially for client/serverless traffic - where near-identical payloads arrive constantly. The vast majority of records are duplicates - that add no information. - -### Option 2 (recommended): Deduplicate and store one record per configuration - -Store one record per distinct configuration and discard the rest. The vast majority of incoming -payloads are duplicates, so dedup is what keeps this feature affordable. The question is what counts -as a "distinct configuration" — i.e. what we deduplicate by. We propose a **dedup key built from the -SDK-set options hash, with the composite natural key as a fallback**: - -> **SDK-set options hash (primary) + `release` + `environment` + `dist` (fallback).** - -#### The required SDK-set options hash - -SDKs **MUST** set the `options_hash` field on the `sdk_config` payload — a **hash of their options -that is stable across instances sharing the same configuration** (same config → same hash; the hash -changes when the config changes). The server uses it as the primary dedup key, so two instances with -the same release+environment+dist but genuinely different options (different hash) are stored as -**distinct records** — within-release variation and drift are visible instead of being collapsed -into the first-seen config. - -- **The hash is computed off both the serialized `options` block and the `integrations` block**, - each normalized by the same [serialization rules](#options-serialization-rules). Since integrations - no longer live in `options`, they must be folded into the hash explicitly and in the same way as - the options — so a change to which integrations are registered, to an integration's own options, or - to its `applied` status yields a different hash. This makes it deterministic and - stable: it is **identical across instances that share a configuration** and **differs when the - configuration differs**. It does not need to be comparable _across_ SDKs, so the specific hash - algorithm is up to each SDK. -- **The hash is computed SDK-side, by design.** Otherwise, we cannot reliably relate events to their - respective config. -- **The hash MUST also be attached to every event the SDK produces**, so events can be - correlated back to the exact config that produced them: - - as an **attribute** on spans, logs, and other attribute-carrying items. We propose `sentry.config_hash` as a semantic attribute. - - as a **context field** on error and transaction events. We propose `sdk_config.hash` as a new context with a single field for now. - -#### The composite natural key (`release` + `environment` + `dist`) — fallback - -Although the hash is required on the wire, a matching stored `sdk_config` record is **not -guaranteed** to exist for a given event's hash (see the caveats below — the config may have been -dropped, sampled out, or stamped before it settled). For those cases we keep the composite natural -key as a fallback lookup. Release alone is **not** sufficient — configuration can legitimately -differ across environments and builds within the same release (per-environment options, -per-deployment overrides, env-var-driven values) — so the fallback key is -`release` + `environment` + `dist`. - -- It is **bounded and predictable** (roughly one record per tuple), **human-readable and directly - queryable** ("show the config for release X in production"), **cheap** (a few stable top-level - fields, no whole-payload processing), and every event already carries these fields. -- Its weakness is `release`: `environment` effectively always has a value (it defaults to - `production`), `dist` is normally absent, so `release` is the load-bearing part — and it has **no - default** and is frequently unset. When it is, the key degrades to roughly `environment` alone - (almost always `production`), collapsing distinct configs into one bucket. It is also blind to - differences _within_ a tuple: instances sharing release+environment+dist but differing in some - option resolve to one (first-seen) record, so genuine variation is not distinguishable at this - level. This is exactly why the hash is the primary key and this is only the fallback. - -#### Correlating an event to its config - -This falls directly out of the dedup key: - -- **By hash (primary):** correlation is **exact** — the event's stamped hash matches exactly one - stored `sdk_config` record. This is the whole reason the hash must also live on events. -- **By composite key (fallback):** when no stored config matches the event's hash — or the event - carries no hash yet (see Timing below) — correlate via the composite key the event already carries - (`release` + `environment` + `dist`). This is free and needs no event changes, but is only - **bucket-level** — if config varied within the tuple, the event resolves to the bucket (a single - first-seen record, or the set of stored variants), not necessarily the exact config that produced - it. - -#### Trade-offs and caveats of the hash - -- **Per-event overhead.** Stamping the hash adds a field to every event, at full event volume. -- **Timing.** The config is not fully settled the instant `init()` returns (debounce window, - `applied` known late). The SDK must ensure the hash it stamps on events matches the config it - reports — e.g. by hashing only config that is stable from `init()`, or by not stamping until the - config has settled. Events emitted before that point may carry no hash (falling back to - bucket-level correlation). -- **Sampling gaps.** For client/serverless SDKs we _may_ sample `sdk_config` (see Sending), so an - event's hash can reference a config that was **never stored** — correlation then fails for exactly - the instances we chose not to persist. - -# Open Questions - -- **What does standing up a new envelope item type actually require?** Introducing `sdk_config` - is not just an SDK + storage change; it needs first-class handling in the ingest pipeline, and - we need to enumerate what that entails before committing. At least: - - **Relay / ingest support.** Registering the new item type, routing it to its own handler, - and any validation/normalization (the `normalized_options` derivation lands here). - - **Rate limiting.** Does `sdk_config` need its own rate-limit category, or does it share an - existing one? How does rate limiting interact with the send cadence (debounced server sends, - client sampling)? What is communicated back to the SDK (e.g. `429` / `Retry-After`) and how - does the SDK back off? - - **Payload / size limits.** What per-item and per-envelope size limits apply, and what happens - when a config exceeds them — reject, or truncate (and if so, which flattened `options` keys to - drop)? - - **Data category, quota & billing.** Which data category does it map to for - quotas/outcomes/billing? The intent is that this is supplementary telemetry, so it most - likely should **not** be billed like events — but that needs to be decided explicitly. - - **Outcomes / observability.** How are dropped or rejected `sdk_config` items recorded - (outcomes, reasons) so we can see ingestion health for this new type? -- **Finalizing the dedup key.** (See [Storing](#storing) above.) We recommend a required SDK-set - options hash as the primary key, with `release` + `environment` + `dist` as the fallback lookup, - but the exact field set, the fallback behavior when no stored config matches an event's hash, and - what SDKs should hash (and how they keep it stable) still need to be nailed down. +The key fields (see [Appendix A](#appendix-a-payload-details) for all fields and exact rules): + +- **`options`:** the **effective** configuration that the SDK runs with (after defaults, environment + variables, and derived values), under native option names. Values are reduced to JSON primitives: + nested objects become dot-notation keys, callbacks become `"[Function]"`, and integrations move to + the `integrations` block. +- **`options_set_by_user`:** the keys that the user explicitly set in `init()`, to tell actual usage + apart from defaults. +- **`integrations`:** the serialized options of each registered integration, plus an optional + `applied` flag that records whether the integration took effect at runtime. For example, the Node + SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. +- **`options_hash`:** a required hash of the configuration (see [Storage](#storage)). +- **`normalized_options`:** a small catalog of options under canonical cross-SDK names (JS + `tracesSampleRate` → `traces_sample_rate`), derived by Relay in the spirit of + [0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md). + +The design keeps SDKs simple: they serialize their existing options object (e.g. JS +`client.getOptions()`), diff its keys against the `init()` argument, and ship no name mapping. New options are captured +automatically, one Relay implementation avoids inconsistent mappings across SDKs, and the catalog can +change, including for stored data, without SDK releases. Sending two option trees or wrapping every +value as `{ value, source }` would need more SDK bookkeeping. The cost is that converted values lose +their original form: for example, we cannot tell whether `integrations` was passed as a function (in +JS, only `integrations` and `stackParser` are affected). + +Sensitive data is scrubbed primarily server-side. SDKs MAY also scrub values that they know to be +sensitive, but they do not guarantee fully scrubbed data. + +## Sending + +**Server SDKs** (Node, Python, Java, Go, …) send one payload per `init()` once the configuration has +settled, because values such as `applied` are only known after `init()` returns. Each SDK MAY choose +how to wait, for example with a short debounce (about 2 seconds) or a lifecycle hook. Short-lived +processes flush the payload on shutdown, and long-lived processes MAY re-send it at a slow interval +(e.g. hourly) in case a send was lost. [Appendix B](#appendix-b-sending-and-storage-details) has the +details. + +**Client SDKs** (browser, mobile, desktop, gaming, …) call `init()` on every page load, app launch, +or session, so sending from every instance produces far more data. We propose exploring: + +| Option | Pros | Cons | +| ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | +| **I: Send on every `init()`** | Simple, same as server SDKs; every instance is represented | Much more data to ingest; an extra request per `init()` | +| **II: Sample** a fraction of `init()` calls (e.g. 1%), assuming configuration is essentially constant per release | Much less overhead | Low-traffic releases and rare configurations may be missed; more configuration surface; harder to reason about and test | +| **III: Attach to the first outgoing envelope**; combinable with I or II, also usable by server SDKs | No extra request; only instances that produce data report | Harder to implement; the first envelope may leave before the configuration settles; instances that never send, e.g. misconfigured ones, never report | + +**Serverless** functions run server SDKs but have one `init()` per short-lived invocation, like +clients, so client strategies such as sampling COULD apply to them. + +## Storage + +Storing every payload (**Option 1**) keeps a complete history, but mostly stores duplicates at a very +high cost. We recommend **Option 2: store one record per distinct configuration**, deduplicated by +**`options_hash` (primary) + `release` + `environment` + `dist` (fallback)**. + +SDKs MUST compute `options_hash` from the serialized `options` and `integrations` blocks, and MUST +attach it to every event: as the proposed `sentry.config_hash` attribute on spans, logs, and other +items with attributes, and in a new `sdk_config.hash` context field on errors and transactions. The +hash links each event to its exact configuration and keeps configurations that differ within one +release separate. + +If no stored record matches an event's hash (e.g. because the payload was lost or sampled out), or +the event has no hash, correlation falls back to `release` + `environment` + `dist`, which every +event already carries. The fallback resolves only to the records of that combination and is coarse +when `release` is unset. # Drawbacks -- **Risk of leaking sensitive data.** Configuration can contain secrets and PII — DSNs, auth - tokens or custom headers in transport options, tunnel URLs, and user-meaningful values in - options like `denyUrls`, `initialScope`, or `serverName`. Even with primitives-only - normalization, we are shipping user configuration to Sentry, and redaction is imperfect. Any - new field an SDK captures is a potential leak, so the capture surface must be curated - carefully. -- **Ingestion and storage overhead.** This is a brand-new payload type sent at potentially very - high volume (every init for client/serverless traffic). Even with sampling and dedup, it adds - ingestion load, a new storage model, and server-side processing that did not exist before. -- **Data may be incomplete or misleading.** The techniques that keep volume down also reduce - fidelity: client sampling can under-sample or miss low-traffic releases and rare - configurations, and when an event has to fall back to the composite key (no stored config matches - its hash), it resolves only to the first-seen config per `release`+`environment`+`dist` bucket even - when config actually varies within it. Decisions made on this data (e.g. deprecating an option that - "looks unused") could be based on a non-representative picture. -- **Lossy representation of runtime options.** Reducing callbacks to `"[Function]"` tells us a - `beforeSend`/`tracesSampler` exists but nothing about what it does. For audit-style use cases - ("warn about confusing behavior") this is a hard limit — we can see that filtering is - configured, not what it filters. -- **Ongoing maintenance burden.** The normalized schema (and any cross-SDK naming catalog) must - be kept in sync with evolving options and integrations across every SDK and language. New - options are invisible until each SDK is updated to capture them, so the dataset always lags - the SDKs, and consistency across SDKs takes continuous effort. -- **Added SDK complexity and runtime cost.** Every SDK gains new machinery: debounced sending, - flush-on-shutdown, opt-in per-integration `applied` tracking, normalization, and (for clients) - sampling. This is more code, more surface for bugs, and some runtime overhead on every init. -- **Unresolved dedup identity.** Storage relies on a good key to deduplicate by. The required options - hash is only as good as each SDK's hashing, and the composite fallback key - (`release`+`environment`+`dist`) is coarse when `release` is absent. Getting this wrong means - either storing too much or collapsing genuinely different configurations together (see Storing). - -# Not in scope / Follow-up work - -This RFC focuses on capturing, transporting, and storing SDK configuration. The following are -explicitly out of scope here and left as follow-up work: - -- **Actually using this data in product.** The downstream use cases from the Motivation - (analytics, setup audits/warnings, showing data sources, configuration-over-time views, a UI - to change configuration, etc.) are not designed here — they build on top of the data this RFC - makes available. -- **Removing SDK metadata from error/transaction events.** Once `sdk_config` is the canonical - home for SDK identity/integration metadata, it makes sense to stop duplicating it on every - error/transaction event. This is not free, though: that metadata is currently searchable and - used on events, so removing it requires a mechanism to backfill it onto events from the stored - `sdk_config` (so events remain searchable/filterable by SDK, version, integrations, etc.). - Designing that backfill is follow-up work. - +- **Sensitive data:** configuration can contain secrets and PII (DSNs, tokens or headers in transport + options, tunnel URLs, `denyUrls`, `initialScope`, `serverName`), and redaction is imperfect. +- **Overhead:** a new, potentially high-volume item (every `init()` for client and serverless + traffic), a new storage model, and new server-side processing. +- **Incomplete or misleading data:** sampling and fallback correlation can give a non-representative + picture, for example when deciding whether an option "looks unused". +- **Lossy callbacks:** `"[Function]"` shows that filtering is configured, not what it filters. +- **Maintenance:** the normalization catalog must track options across all SDKs. +- **SDK complexity:** delayed sending, flushing, `applied` tracking, serialization, and sampling add + code and runtime cost on every `init()`. +- **Deduplication identity:** a poor hash or a missing `release` either stores too much or merges + distinct configurations. + +# Unresolved questions + +- **Ingest requirements for a new item type:** Relay support (routing, validation, deriving + `normalized_options`), rate limiting (own or shared category, interaction with the send cadence, + `429`/`Retry-After` and SDK backoff), size limits (reject or truncate), data category, quota + consumption, and billing (likely not billed like events), and outcomes for dropped items. +- **Deduplication key:** the exact fields, the fallback behavior, and what SDKs hash and how they keep + the hash stable. +- **EAP storage:** EAP prefers a single top-level `attributes` field, which would mean storing blocks + such as `options` as JSON-valued attributes that remain efficiently queryable. Whether to use EAP is + decided with the storage design. +- **Item type name** and **client SDK send strategy** (Options I to III). + +## Out of scope + +- Product features built on this data (analytics, audits, data source views, configuration history, + configuration UI). +- Removing SDK metadata from events. Events are searchable by it, so this first requires backfilling + it onto events from stored `sdk_config` records. + +# Appendix A: Payload details + +**Fields** + +- `version`: integer payload schema version, starting at `1` and independent of `sdk.version`. It is + bumped only for changes that consumers must know about, so the server can select the right parser + instead of guessing from the field set. +- `timestamp`: when the SDK sent the payload, used to order records and track changes. ISO 8601 or + epoch seconds, following Sentry conventions. +- `sdk`: the same data as on events today; `sdk_config` becomes its canonical place. The `integrations` + block captures integration information in more detail than the event's list of integration names; + events keep their existing metadata for now (see [Out of scope](#out-of-scope)). +- `meta`: `release`, `environment`, `dist`, and runtime information, regardless of how they were set. +- `options_set_by_user`: uses the flattened keys of `options`; each key corresponds 1:1 to a key in + `options`. +- `integrations`: each integration's options are nested under `options`, so they cannot collide with + status keys. `applied` is `true` if the integration took effect (e.g. patched Express), `false` if + not (e.g. its framework is absent), and omitted if unknown. It is opt-in, added where the signal is + useful, such as framework integrations. +- `normalized_options`: entries are `{ "key": , "value": }`; whether + the user set an option is checked in `options_set_by_user`. +- `_other`: free-form, SDK-specific data; useful keys can later become first-class fields. + +**Serialization rules** for `options` and each integration's `options`, which make the output +deterministic and consistent across SDKs: + +1. Primitives are sent unchanged. Arrays stay arrays, with each element serialized by these same + rules. +2. Nested objects are flattened: `dataCollection: { http: { bodies: true } }` becomes + `"dataCollection.http.bodies": true`. +3. Functions become `"[Function]"`. SDKs MAY include the name (`"[Function: beforeSend]"`), but + consumers must not rely on it. +4. The `integrations` option is omitted; its data is in the `integrations` block. +5. Other non-serializable values become type markers such as `"[SomeType]"`, following the SDK's + existing normalization convention. + +**Initial normalization catalog**, maintained alongside Relay and expected to grow. `release`, +`environment`, and `dist` are not included, because they are already in `meta`. + +| Canonical key | Meaning | Native names (JS → Python) | +| ------------------------- | --------------------------- | --------------------------------------------------- | +| `sample_rate` | Error sample rate | `sampleRate` → `sample_rate` | +| `traces_sample_rate` | Tracing sample rate | `tracesSampleRate` → `traces_sample_rate` | +| `profiles_sample_rate` | Profiling sample rate | `profilesSampleRate` → `profiles_sample_rate` | +| `send_default_pii` | Whether default PII is sent | `sendDefaultPii` → `send_default_pii` | +| `debug` | Debug logging enabled | `debug` → `debug` | +| `before_send` | Whether the hook is set | `beforeSend` → `before_send` | +| `before_send_transaction` | Whether the hook is set | `beforeSendTransaction` → `before_send_transaction` | +| `before_send_span` | Whether the hook is set | `beforeSendSpan` → `before_send_span` | +| `before_send_log` | Whether the hook is set | `beforeSendLog` → `before_send_log` | +| `ignore_spans` | Span-ignore rules | `ignoreSpans` → `ignore_spans` | +| `traces_sampler` | Whether the hook is set | `tracesSampler` → `traces_sampler` | +| `data_collection.*` | All flattened sub-keys | `dataCollection.*` → `data_collection.*` | + +# Appendix B: Sending and storage details + +- **Waiting for settled configuration:** values known only after `init()` include `applied`, lazily + registered integrations, and an asynchronously detected release or environment. The debounce MAY be + configurable for setups that settle later. A lifecycle hook that guarantees settled options (e.g. + after the first request, a post-init hook, or the first idle event loop) is preferable. The goal is + one stable payload, not a stream of updates; later configuration changes are a separate concern. +- **Short-lived processes** (serverless, CLI) flush wherever they already flush events, and MAY reuse + their client report flushing. If a process dies first, an equivalent instance will report. +- **Periodic re-send:** if Relay cannot guarantee that a `200` response means the payload was + persisted, a single lost send would leave a long-running server without stored configuration. +- **Hash:** identical configurations produce identical hashes, and any change to the options, the + registered integrations, their options, or `applied` changes the hash. Algorithms do not need to + match across SDKs. The hash is computed SDK-side so that events can carry it. +- **Hash caveats:** it adds a field to every event. Because configuration settles after `init()`, the + hash on events must match the reported configuration (e.g. by hashing only values that are stable + from `init()`, or by stamping events only after settling), so early events may lack it. Payloads that + were sampled out leave hashes without a stored record. +- **Fallback key:** `release` alone is insufficient, because configuration can differ by environment + and build. The composite key is bounded, human-readable, and cheap, but `environment` defaults to + `production` and `dist` is usually absent, so the key depends on `release`, which is often unset. It + resolves to the first-seen record or to the set of stored variants. From acc3573f25e53b64cbbfd1e0749948a9ddbcd2ce Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Fri, 25 Sep 2026 14:43:26 +0200 Subject: [PATCH 32/51] small clarifications --- text/0162-capture-sdk-options.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 750cad0f..81446f82 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -139,7 +139,7 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for all fields and - **`options`:** the **effective** configuration that the SDK runs with (after defaults, environment variables, and derived values), under native option names. Values are reduced to JSON primitives: nested objects become dot-notation keys, callbacks become `"[Function]"`, and integrations move to - the `integrations` block. + the `integrations` block. Option names should reflect the names a user would use to set the values. - **`options_set_by_user`:** the keys that the user explicitly set in `init()`, to tell actual usage apart from defaults. - **`integrations`:** the serialized options of each registered integration, plus an optional @@ -249,10 +249,11 @@ when `release` is unset. - `meta`: `release`, `environment`, `dist`, and runtime information, regardless of how they were set. - `options_set_by_user`: uses the flattened keys of `options`; each key corresponds 1:1 to a key in `options`. -- `integrations`: each integration's options are nested under `options`, so they cannot collide with +- `integrations`: each integration's options are optionally nested under `options`, so they cannot collide with status keys. `applied` is `true` if the integration took effect (e.g. patched Express), `false` if not (e.g. its framework is absent), and omitted if unknown. It is opt-in, added where the signal is - useful, such as framework integrations. + useful, such as framework integrations. `options` MAY also be omitted if no options are sent, + if they are hard to access in a given SDK, or if they are low-value. - `normalized_options`: entries are `{ "key": , "value": }`; whether the user set an option is checked in `options_set_by_user`. - `_other`: free-form, SDK-specific data; useful keys can later become first-class fields. @@ -262,6 +263,7 @@ deterministic and consistent across SDKs: 1. Primitives are sent unchanged. Arrays stay arrays, with each element serialized by these same rules. + a. An SDK MAY normalize specific options if it makes sense, e.g. stripping out user-specific paths or similar. 2. Nested objects are flattened: `dataCollection: { http: { bodies: true } }` becomes `"dataCollection.http.bodies": true`. 3. Functions become `"[Function]"`. SDKs MAY include the name (`"[Function: beforeSend]"`), but From b3fa575b6e0d4334666a1b65f52d9fe4e47e1f30 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Tue, 29 Sep 2026 09:22:35 +0200 Subject: [PATCH 33/51] mention options are best-effort --- text/0162-capture-sdk-options.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 81446f82..5b4a5e3b 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -141,7 +141,8 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for all fields and nested objects become dot-notation keys, callbacks become `"[Function]"`, and integrations move to the `integrations` block. Option names should reflect the names a user would use to set the values. - **`options_set_by_user`:** the keys that the user explicitly set in `init()`, to tell actual usage - apart from defaults. + apart from defaults. This should be best-effort - it MAY be incomplete if users add configuration + in alternate paths or similar. - **`integrations`:** the serialized options of each registered integration, plus an optional `applied` flag that records whether the integration took effect at runtime. For example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. From 37963f944e1173fd691e6b8f430a6be39b3a303b Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 08:40:42 +0200 Subject: [PATCH 34/51] better guard --- text/0162-capture-sdk-options.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 5b4a5e3b..e6e54e69 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -142,7 +142,8 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for all fields and the `integrations` block. Option names should reflect the names a user would use to set the values. - **`options_set_by_user`:** the keys that the user explicitly set in `init()`, to tell actual usage apart from defaults. This should be best-effort - it MAY be incomplete if users add configuration - in alternate paths or similar. + in alternate paths or similar. If it is not possible to enumerate options automatically, SDKs MAY + send a hand-picked subset of options here only. - **`integrations`:** the serialized options of each registered integration, plus an optional `applied` flag that records whether the integration took effect at runtime. For example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. @@ -312,4 +313,4 @@ deterministic and consistent across SDKs: - **Fallback key:** `release` alone is insufficient, because configuration can differ by environment and build. The composite key is bounded, human-readable, and cheap, but `environment` defaults to `production` and `dist` is usually absent, so the key depends on `release`, which is often unset. It - resolves to the first-seen record or to the set of stored variants. + resolves to the first-seen record or to the set of stored variants. \ No newline at end of file From cc33e41176fe8f1e4883e6b7ccc178860ae14553 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 10:18:03 +0200 Subject: [PATCH 35/51] move to eap --- text/0162-capture-sdk-options.md | 243 +++++++++++++++++-------------- 1 file changed, 131 insertions(+), 112 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index e6e54e69..26e4ce91 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -3,7 +3,7 @@ - RFC PR: https://github.com/getsentry/rfcs/pull/162 - RFC Status: draft - RFC Author: @mydea -- RFC Approver: +- RFC Approver: # Summary @@ -67,90 +67,105 @@ one language-agnostic schema for all SDKs, starting with JavaScript. Older Relay versions discard unknown item types without affecting the rest of the envelope, so SDKs can send `sdk_config` unconditionally. -SDKs send every field in this example except `normalized_options`, which Relay adds at ingestion: +The item follows the EAP format, so it can be stored in EAP as a new +trace item type: a container with `items`, each with a `timestamp` and typed `attributes`. Existing +[Sentry conventions](https://getsentry.github.io/sentry-conventions/attributes/) are reused where +they fit; everything else lives under `sentry.sdk_config.*` in dot notation. SDKs send every +attribute in this example except `sentry.sdk_config.normalized.*`, which Relay adds at ingestion: ```json +{"type":"sdk_config","item_count":1,"content_type":"application/vnd.sentry.items.sdk-config+json"} { - "version": 1, - "timestamp": "2026-09-21T12:00:00Z", - - "sdk": { - "name": "sentry.javascript.node", - "version": "10.0.0", - "packages": [{ "name": "npm:@sentry/node", "version": "10.0.0" }] - }, - - "meta": { - "release": "my-app@1.2.3", - "environment": "production", - "dist": "42", - "runtime": { "name": "node", "version": "20.11.0" } - }, - - "options": { - "dsn": "https://@o0.ingest.sentry.io/0", - "sampleRate": 1.0, - "tracesSampleRate": 0.2, - "sendDefaultPii": true, - "debug": false, - "environment": "production", - "beforeSend": "[Function]", - "tracesSampler": "[Function]", - "denyUrls": ["https://example.com/ignore"], - "dataCollection.http.bodies": true, - "dataCollection.http.headers": false - }, - - "options_set_by_user": [ - "dsn", - "tracesSampleRate", - "sendDefaultPii", - "beforeSend", - "tracesSampler", - "denyUrls" - ], - - "options_hash": "9f2c1a7e", - - "normalized_options": { - "sample_rate": { "key": "sampleRate", "value": 1.0 }, - "traces_sample_rate": { "key": "tracesSampleRate", "value": 0.2 }, - "send_default_pii": { "key": "sendDefaultPii", "value": true }, - "debug": { "key": "debug", "value": false }, - "before_send": { "key": "beforeSend", "value": "[Function]" } - }, - - "integrations": { - "InboundFilters": { "options": {} }, - "ExpressIntegration": { "applied": true, "options": {} }, - "FastifyIntegration": { "applied": false, "options": {} }, - "KoaIntegration": { "options": {} }, - "MyIntegration": { - "options": { "filter": "aaa", "shouldLog": "[Function]" } + "items": [ + { + "timestamp": 1790000000.0, + "attributes": { + "sentry.sdk.name": { "type": "string", "value": "sentry.javascript.node" }, + "sentry.sdk.version": { "type": "string", "value": "10.0.0" }, + "sentry.sdk.packages": { "type": "array", "value": ["npm:@sentry/node@10.0.0"] }, + "sentry.sdk.integrations": { + "type": "array", + "value": ["InboundFilters", "Express", "Fastify", "Koa", "MyIntegration"] + }, + "sentry.release": { "type": "string", "value": "my-app@1.2.3" }, + "sentry.environment": { "type": "string", "value": "production" }, + "sentry.dist": { "type": "string", "value": "42" }, + "process.runtime.name": { "type": "string", "value": "node" }, + "process.runtime.version": { "type": "string", "value": "20.11.0" }, + + "sentry.sdk_config.version": { "type": "integer", "value": 1 }, + "sentry.sdk_config.hash": { "type": "string", "value": "9f2c1a7e" }, + + "sentry.sdk_config.option.dsn": { "type": "string", "value": "https://@o0.ingest.sentry.io/0" }, + "sentry.sdk_config.option.sampleRate": { "type": "double", "value": 1.0 }, + "sentry.sdk_config.option.tracesSampleRate": { "type": "double", "value": 0.2 }, + "sentry.sdk_config.option.sendDefaultPii": { "type": "boolean", "value": true }, + "sentry.sdk_config.option.beforeSend": { "type": "string", "value": "[Function]" }, + "sentry.sdk_config.option.denyUrls": { "type": "array", "value": ["https://example.com/ignore"] }, + "sentry.sdk_config.option.dataCollection.http.bodies": { "type": "boolean", "value": true }, + "sentry.sdk_config.options_set_by_user": { + "type": "array", + "value": ["dsn", "tracesSampleRate", "sendDefaultPii", "beforeSend", "denyUrls"] + }, + + "sentry.sdk_config.integration.Express.applied": { "type": "boolean", "value": true }, + "sentry.sdk_config.integration.Fastify.applied": { "type": "boolean", "value": false }, + "sentry.sdk_config.integration.MyIntegration.option.filter": { "type": "string", "value": "aaa" }, + "sentry.sdk_config.integration.MyIntegration.option.shouldLog": { "type": "string", "value": "[Function]" }, + + "sentry.sdk_config.normalized.sample_rate": { "type": "double", "value": 1.0 }, + "sentry.sdk_config.normalized.sample_rate.original": { "type": "string", "value": "sampleRate" }, + "sentry.sdk_config.normalized.traces_sample_rate": { "type": "double", "value": 0.2 }, + "sentry.sdk_config.normalized.traces_sample_rate.original": { "type": "string", "value": "tracesSampleRate" }, + "sentry.sdk_config.normalized.send_default_pii": { "type": "boolean", "value": true }, + "sentry.sdk_config.normalized.send_default_pii.original": { "type": "string", "value": "sendDefaultPii" }, + "sentry.sdk_config.normalized.before_send": { "type": "string", "value": "[Function]" }, + "sentry.sdk_config.normalized.before_send.original": { "type": "string", "value": "beforeSend" } + } } - }, - - "_other": {} + ] } ``` The key fields (see [Appendix A](#appendix-a-payload-details) for all fields and exact rules): -- **`options`:** the **effective** configuration that the SDK runs with (after defaults, environment - variables, and derived values), under native option names. Values are reduced to JSON primitives: - nested objects become dot-notation keys, callbacks become `"[Function]"`, and integrations move to - the `integrations` block. Option names should reflect the names a user would use to set the values. -- **`options_set_by_user`:** the keys that the user explicitly set in `init()`, to tell actual usage - apart from defaults. This should be best-effort - it MAY be incomplete if users add configuration - in alternate paths or similar. If it is not possible to enumerate options automatically, SDKs MAY - send a hand-picked subset of options here only. -- **`integrations`:** the serialized options of each registered integration, plus an optional - `applied` flag that records whether the integration took effect at runtime. For example, the Node - SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. -- **`options_hash`:** a required hash of the configuration (see [Storage](#storage)). -- **`normalized_options`:** a small catalog of options under canonical cross-SDK names (JS - `tracesSampleRate` → `traces_sample_rate`), derived by Relay in the spirit of - [0116-sentry-semantic-conventions](./0116-sentry-semantic-conventions.md). +- **`sentry.sdk_config.option.`:** the **effective** configuration that the SDK runs with (after + defaults, environment variables, and derived values), under native option names. Values are reduced + to attribute types: nested objects become dot-notation keys, callbacks become `"[Function]"`, and + integrations are covered by the integration attributes. Option names should reflect the names a + user would use to set the values. +- **`sentry.sdk_config.options_set_by_user`:** the option keys that the user explicitly set in + `init()`, to tell actual usage apart from defaults. This should be best-effort - it MAY be + incomplete if users add configuration in alternate paths or similar. If it is not possible to + enumerate options automatically, SDKs MAY send a hand-picked subset of options here only. +- **Integrations:** `sentry.sdk.integrations` lists every registered integration. Per integration, + `sentry.sdk_config.integration..option.` holds its serialized options, and an optional + `sentry.sdk_config.integration..applied` records whether it took effect at runtime. For + example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically + uses only one. +- **`sentry.sdk_config.hash`:** a required hash of the configuration (see [Storage](#storage)). +- **`sentry.sdk_config.normalized.`:** a small catalog of options under canonical cross-SDK + names (JS `tracesSampleRate` → `traces_sample_rate`), derived by Relay. +- **`sentry.sdk_config.normalized..original`:** the name that this property is called in the `options` (JS `tracesSampleRate`) + to allow display of the value in a way that makes sense for a user. + +New attributes to add to Sentry conventions: + +| Attribute | Type | Example | +| --------------------------------------------------- | -------- | ---------------------------------------------------------------- | +| `sentry.sdk.packages` | string[] | `["npm:@sentry/node@10.0.0"]` | +| `sentry.sdk_config.version` | integer | `1` | +| `sentry.sdk_config.hash` | string | `"9f2c1a7e"` | +| `sentry.sdk_config.option.` | any | `sentry.sdk_config.option.sampleRate=1.0` | +| `sentry.sdk_config.options_set_by_user` | string[] | `["dsn", "tracesSampleRate"]` | +| `sentry.sdk_config.integration..applied` | boolean | `true` | +| `sentry.sdk_config.integration..option.` | any | `...MyIntegration.option.filter="aaa"` | +| `sentry.sdk_config.normalized.` | any | `...normalized.traces_sample_rate=0.2` | +| `sentry.sdk_config.normalized..original` | string | `sentry.sdk_config.normalized.sample_rate.original="sampleRate"` | +| `sentry.sdk_config.other.` | any | `sentry.sdk_config.other.foo="bar"` | + +Reused as-is: `sentry.sdk.name`, `sentry.sdk.version`, `sentry.sdk.integrations`, `sentry.release`, +`sentry.environment`, `sentry.dist`, `process.runtime.name`, `process.runtime.version`. The design keeps SDKs simple: they serialize their existing options object (e.g. JS `client.getOptions()`), diff its keys against the `init()` argument, and ship no name mapping. New options are captured @@ -188,11 +203,11 @@ clients, so client strategies such as sampling COULD apply to them. Storing every payload (**Option 1**) keeps a complete history, but mostly stores duplicates at a very high cost. We recommend **Option 2: store one record per distinct configuration**, deduplicated by -**`options_hash` (primary) + `release` + `environment` + `dist` (fallback)**. +**`sentry.sdk_config.hash` (primary) + `release` + `environment` + `dist` (fallback)**. -SDKs MUST compute `options_hash` from the serialized `options` and `integrations` blocks, and MUST -attach it to every event: as the proposed `sentry.config_hash` attribute on spans, logs, and other -items with attributes, and in a new `sdk_config.hash` context field on errors and transactions. The +SDKs MUST compute the hash from the serialized option and integration attributes, and MUST attach it +to every event: as the same `sentry.sdk_config.hash` attribute on spans, logs, and other items with +attributes, and in a new `sdk_config.hash` context field on errors and transactions. The hash links each event to its exact configuration and keeps configurations that differ within one release separate. @@ -219,14 +234,15 @@ when `release` is unset. # Unresolved questions - **Ingest requirements for a new item type:** Relay support (routing, validation, deriving - `normalized_options`), rate limiting (own or shared category, interaction with the send cadence, + normalized attributes), rate limiting (own or shared category, interaction with the send cadence, `429`/`Retry-After` and SDK backoff), size limits (reject or truncate), data category, quota consumption, and billing (likely not billed like events), and outcomes for dropped items. - **Deduplication key:** the exact fields, the fallback behavior, and what SDKs hash and how they keep the hash stable. -- **EAP storage:** EAP prefers a single top-level `attributes` field, which would mean storing blocks - such as `options` as JSON-valued attributes that remain efficiently queryable. Whether to use EAP is - decided with the storage design. +- **EAP storage:** a new `TraceItemType` (sentry-protos, Snuba, Relay, Sentry search). EAP items + expire after their retention period, so long-running processes must re-send within it. EAP could + deduplicate via a deterministic `item_id` from the hash plus a bucketed `timestamp` (as preprod + does), to be confirmed with the EAP team. Limits on attribute count and size per item. - **Item type name** and **client SDK send strategy** (Options I to III). ## Out of scope @@ -240,42 +256,45 @@ when `release` is unset. **Fields** -- `version`: integer payload schema version, starting at `1` and independent of `sdk.version`. It is - bumped only for changes that consumers must know about, so the server can select the right parser - instead of guessing from the field set. -- `timestamp`: when the SDK sent the payload, used to order records and track changes. ISO 8601 or - epoch seconds, following Sentry conventions. -- `sdk`: the same data as on events today; `sdk_config` becomes its canonical place. The `integrations` - block captures integration information in more detail than the event's list of integration names; - events keep their existing metadata for now (see [Out of scope](#out-of-scope)). -- `meta`: `release`, `environment`, `dist`, and runtime information, regardless of how they were set. -- `options_set_by_user`: uses the flattened keys of `options`; each key corresponds 1:1 to a key in - `options`. -- `integrations`: each integration's options are optionally nested under `options`, so they cannot collide with - status keys. `applied` is `true` if the integration took effect (e.g. patched Express), `false` if - not (e.g. its framework is absent), and omitted if unknown. It is opt-in, added where the signal is - useful, such as framework integrations. `options` MAY also be omitted if no options are sent, - if they are hard to access in a given SDK, or if they are low-value. -- `normalized_options`: entries are `{ "key": , "value": }`; whether - the user set an option is checked in `options_set_by_user`. -- `_other`: free-form, SDK-specific data; useful keys can later become first-class fields. - -**Serialization rules** for `options` and each integration's `options`, which make the output +- `timestamp`: top-level item field; when the SDK sent the payload, in epoch seconds, used to order + records and track changes. +- `trace_id`: EAP requires one, but a configuration belongs to no trace. SDKs omit it, and Relay + derives a stable one (e.g. from the hash). +- `sentry.sdk_config.version`: integer payload schema version, starting at `1` and independent of + `sentry.sdk.version`. It is bumped only for changes that consumers must know about, so the server + can select the right parser instead of guessing from the attribute set. +- `sentry.sdk.*`: the same data as the `sdk` object on events today; `sdk_config` becomes its + canonical place. Events keep their existing metadata for now (see [Out of scope](#out-of-scope)). +- `sentry.release`, `sentry.environment`, `sentry.dist`, `process.runtime.*`: the effective values, + regardless of how they were set. +- `sentry.sdk_config.options_set_by_user`: each entry corresponds 1:1 to a + `sentry.sdk_config.option.` attribute. +- `sentry.sdk_config.integration..applied`: `true` if the integration took effect (e.g. patched + Express), `false` if not (e.g. its framework is absent), and omitted if unknown. It is opt-in, + added where the signal is useful, such as framework integrations. Integration options MAY be + omitted if there are none, if they are hard to access in a given SDK, or if they are low-value. +- `sentry.sdk_config.normalized.`: holds the effective value; the native name is known from the + catalog, and whether the user set it is checked in `options_set_by_user`. +- `sentry.sdk_config.other.`: free-form, SDK-specific data; useful keys can later become + first-class attributes. + +**Serialization rules** for option and integration option attributes, which make the output deterministic and consistent across SDKs: -1. Primitives are sent unchanged. Arrays stay arrays, with each element serialized by these same - rules. +1. Primitives are sent unchanged. Arrays of one primitive type stay arrays; other arrays have each + element converted to a string by these same rules. a. An SDK MAY normalize specific options if it makes sense, e.g. stripping out user-specific paths or similar. 2. Nested objects are flattened: `dataCollection: { http: { bodies: true } }` becomes - `"dataCollection.http.bodies": true`. + `sentry.sdk_config.option.dataCollection.http.bodies`. 3. Functions become `"[Function]"`. SDKs MAY include the name (`"[Function: beforeSend]"`), but consumers must not rely on it. -4. The `integrations` option is omitted; its data is in the `integrations` block. -5. Other non-serializable values become type markers such as `"[SomeType]"`, following the SDK's +4. The `integrations` option is omitted; its data is in the integration attributes. +5. `null` and `undefined` values are omitted, since attributes have no null type. +6. Other non-serializable values become type markers such as `"[SomeType]"`, following the SDK's existing normalization convention. **Initial normalization catalog**, maintained alongside Relay and expected to grow. `release`, -`environment`, and `dist` are not included, because they are already in `meta`. +`environment`, and `dist` are not included, because they already have their own attributes. | Canonical key | Meaning | Native names (JS → Python) | | ------------------------- | --------------------------- | --------------------------------------------------- | @@ -313,4 +332,4 @@ deterministic and consistent across SDKs: - **Fallback key:** `release` alone is insufficient, because configuration can differ by environment and build. The composite key is bounded, human-readable, and cheap, but `environment` defaults to `production` and `dist` is usually absent, so the key depends on `release`, which is often unset. It - resolves to the first-seen record or to the set of stored variants. \ No newline at end of file + resolves to the first-seen record or to the set of stored variants. From bf60e53cf87b0053a9f49aa6a6b5d2d2ec4c5689 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 10:21:14 +0200 Subject: [PATCH 36/51] cleanup --- text/0162-capture-sdk-options.md | 27 ++------------------------- 1 file changed, 2 insertions(+), 25 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 26e4ce91..e5343b47 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -162,7 +162,6 @@ New attributes to add to Sentry conventions: | `sentry.sdk_config.integration..option.` | any | `...MyIntegration.option.filter="aaa"` | | `sentry.sdk_config.normalized.` | any | `...normalized.traces_sample_rate=0.2` | | `sentry.sdk_config.normalized..original` | string | `sentry.sdk_config.normalized.sample_rate.original="sampleRate"` | -| `sentry.sdk_config.other.` | any | `sentry.sdk_config.other.foo="bar"` | Reused as-is: `sentry.sdk.name`, `sentry.sdk.version`, `sentry.sdk.integrations`, `sentry.release`, `sentry.environment`, `sentry.dist`, `process.runtime.name`, `process.runtime.version`. @@ -254,29 +253,7 @@ when `release` is unset. # Appendix A: Payload details -**Fields** - -- `timestamp`: top-level item field; when the SDK sent the payload, in epoch seconds, used to order - records and track changes. -- `trace_id`: EAP requires one, but a configuration belongs to no trace. SDKs omit it, and Relay - derives a stable one (e.g. from the hash). -- `sentry.sdk_config.version`: integer payload schema version, starting at `1` and independent of - `sentry.sdk.version`. It is bumped only for changes that consumers must know about, so the server - can select the right parser instead of guessing from the attribute set. -- `sentry.sdk.*`: the same data as the `sdk` object on events today; `sdk_config` becomes its - canonical place. Events keep their existing metadata for now (see [Out of scope](#out-of-scope)). -- `sentry.release`, `sentry.environment`, `sentry.dist`, `process.runtime.*`: the effective values, - regardless of how they were set. -- `sentry.sdk_config.options_set_by_user`: each entry corresponds 1:1 to a - `sentry.sdk_config.option.` attribute. -- `sentry.sdk_config.integration..applied`: `true` if the integration took effect (e.g. patched - Express), `false` if not (e.g. its framework is absent), and omitted if unknown. It is opt-in, - added where the signal is useful, such as framework integrations. Integration options MAY be - omitted if there are none, if they are hard to access in a given SDK, or if they are low-value. -- `sentry.sdk_config.normalized.`: holds the effective value; the native name is known from the - catalog, and whether the user set it is checked in `options_set_by_user`. -- `sentry.sdk_config.other.`: free-form, SDK-specific data; useful keys can later become - first-class attributes. +For payload fields, see [Payload](#payload). **Serialization rules** for option and integration option attributes, which make the output deterministic and consistent across SDKs: @@ -288,7 +265,7 @@ deterministic and consistent across SDKs: `sentry.sdk_config.option.dataCollection.http.bodies`. 3. Functions become `"[Function]"`. SDKs MAY include the name (`"[Function: beforeSend]"`), but consumers must not rely on it. -4. The `integrations` option is omitted; its data is in the integration attributes. +4. The `integrations` option is omitted; its data is in the `sentry.sdk.integrations` attributes. 5. `null` and `undefined` values are omitted, since attributes have no null type. 6. Other non-serializable values become type markers such as `"[SomeType]"`, following the SDK's existing normalization convention. From 1139b7ed665ba54ca7c0fcf162d1af191aac15e1 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 10:29:37 +0200 Subject: [PATCH 37/51] tweaks, streamlining --- text/0162-capture-sdk-options.md | 37 ++++++++++++++------------------ 1 file changed, 16 insertions(+), 21 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index e5343b47..1638052a 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -179,24 +179,17 @@ sensitive, but they do not guarantee fully scrubbed data. ## Sending -**Server SDKs** (Node, Python, Java, Go, …) send one payload per `init()` once the configuration has +We propose to send one payload per `init()` once the configuration has settled, because values such as `applied` are only known after `init()` returns. Each SDK MAY choose -how to wait, for example with a short debounce (about 2 seconds) or a lifecycle hook. Short-lived -processes flush the payload on shutdown, and long-lived processes MAY re-send it at a slow interval -(e.g. hourly) in case a send was lost. [Appendix B](#appendix-b-sending-and-storage-details) has the +how to wait, for example with a short debounce (e.g. 5 seconds) or a lifecycle hook. Short-lived +processes MAY flush the payload on shutdown. [Appendix B](#appendix-b-sending-and-storage-details) has the details. -**Client SDKs** (browser, mobile, desktop, gaming, …) call `init()` on every page load, app launch, -or session, so sending from every instance produces far more data. We propose exploring: +We propose to do this for all kindes of SDKs - while client SDKs will send more redundant data than server SDKs, +the overall volume will still be low compared to e.g. metrics, spans or logs, and we'll server-side dedupe the records for storage. -| Option | Pros | Cons | -| ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | -| **I: Send on every `init()`** | Simple, same as server SDKs; every instance is represented | Much more data to ingest; an extra request per `init()` | -| **II: Sample** a fraction of `init()` calls (e.g. 1%), assuming configuration is essentially constant per release | Much less overhead | Low-traffic releases and rare configurations may be missed; more configuration surface; harder to reason about and test | -| **III: Attach to the first outgoing envelope**; combinable with I or II, also usable by server SDKs | No extra request; only instances that produce data report | Harder to implement; the first envelope may leave before the configuration settles; instances that never send, e.g. misconfigured ones, never report | - -**Serverless** functions run server SDKs but have one `init()` per short-lived invocation, like -clients, so client strategies such as sampling COULD apply to them. +**Server SDKs** (Node, Python, Java, Go, …) MAY re-send it at a slow interval +(e.g. hourly) in case a send was lost, and to account for data retention dropping old records for very long-lived processes. ## Storage @@ -236,13 +229,10 @@ when `release` is unset. normalized attributes), rate limiting (own or shared category, interaction with the send cadence, `429`/`Retry-After` and SDK backoff), size limits (reject or truncate), data category, quota consumption, and billing (likely not billed like events), and outcomes for dropped items. -- **Deduplication key:** the exact fields, the fallback behavior, and what SDKs hash and how they keep - the hash stable. - **EAP storage:** a new `TraceItemType` (sentry-protos, Snuba, Relay, Sentry search). EAP items expire after their retention period, so long-running processes must re-send within it. EAP could deduplicate via a deterministic `item_id` from the hash plus a bucketed `timestamp` (as preprod does), to be confirmed with the EAP team. Limits on attribute count and size per item. -- **Item type name** and **client SDK send strategy** (Options I to III). ## Out of scope @@ -260,7 +250,7 @@ deterministic and consistent across SDKs: 1. Primitives are sent unchanged. Arrays of one primitive type stay arrays; other arrays have each element converted to a string by these same rules. - a. An SDK MAY normalize specific options if it makes sense, e.g. stripping out user-specific paths or similar. + a. An SDK MAY normalize specific options if it makes sense, e.g. stripping out user-specific paths, unstable options or similar. 2. Nested objects are flattened: `dataCollection: { http: { bodies: true } }` becomes `sentry.sdk_config.option.dataCollection.http.bodies`. 3. Functions become `"[Function]"`. SDKs MAY include the name (`"[Function: beforeSend]"`), but @@ -295,17 +285,22 @@ deterministic and consistent across SDKs: configurable for setups that settle later. A lifecycle hook that guarantees settled options (e.g. after the first request, a post-init hook, or the first idle event loop) is preferable. The goal is one stable payload, not a stream of updates; later configuration changes are a separate concern. + This should be best-effort, it is understood that late config changes MAY not be correctly captured. - **Short-lived processes** (serverless, CLI) flush wherever they already flush events, and MAY reuse their client report flushing. If a process dies first, an equivalent instance will report. -- **Periodic re-send:** if Relay cannot guarantee that a `200` response means the payload was - persisted, a single lost send would leave a long-running server without stored configuration. +- **Periodic re-send:** Relay cannot guarantee that a `200` response means the payload was + persisted. To accomodate this, as well as retention period dropping config after longer time periods, + a single lost send would leave a long-running server without stored configuration. - **Hash:** identical configurations produce identical hashes, and any change to the options, the registered integrations, their options, or `applied` changes the hash. Algorithms do not need to match across SDKs. The hash is computed SDK-side so that events can carry it. - **Hash caveats:** it adds a field to every event. Because configuration settles after `init()`, the hash on events must match the reported configuration (e.g. by hashing only values that are stable from `init()`, or by stamping events only after settling), so early events may lack it. Payloads that - were sampled out leave hashes without a stored record. + were sampled out leave hashes without a stored record. + The current hash should be put on other telemetry items, even if the sdk_config has not been sent yet. + It is understood that this means that the final hash MAY differ from the one attached to early records. + Those may be linked by fallback key only (see below). - **Fallback key:** `release` alone is insufficient, because configuration can differ by environment and build. The composite key is bounded, human-readable, and cheap, but `environment` defaults to `production` and `dist` is usually absent, so the key depends on `release`, which is often unset. It From 5cb3e903183315432f2dbc5736e9b1436dd326db Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 10:33:49 +0200 Subject: [PATCH 38/51] prettier --- text/0162-capture-sdk-options.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 1638052a..55020221 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -189,7 +189,7 @@ We propose to do this for all kindes of SDKs - while client SDKs will send more the overall volume will still be low compared to e.g. metrics, spans or logs, and we'll server-side dedupe the records for storage. **Server SDKs** (Node, Python, Java, Go, …) MAY re-send it at a slow interval -(e.g. hourly) in case a send was lost, and to account for data retention dropping old records for very long-lived processes. +(e.g. hourly) in case a send was lost, and to account for data retention dropping old records for very long-lived processes. ## Storage @@ -289,7 +289,7 @@ deterministic and consistent across SDKs: - **Short-lived processes** (serverless, CLI) flush wherever they already flush events, and MAY reuse their client report flushing. If a process dies first, an equivalent instance will report. - **Periodic re-send:** Relay cannot guarantee that a `200` response means the payload was - persisted. To accomodate this, as well as retention period dropping config after longer time periods, + persisted. To accomodate this, as well as retention period dropping config after longer time periods, a single lost send would leave a long-running server without stored configuration. - **Hash:** identical configurations produce identical hashes, and any change to the options, the registered integrations, their options, or `applied` changes the hash. Algorithms do not need to @@ -297,9 +297,9 @@ deterministic and consistent across SDKs: - **Hash caveats:** it adds a field to every event. Because configuration settles after `init()`, the hash on events must match the reported configuration (e.g. by hashing only values that are stable from `init()`, or by stamping events only after settling), so early events may lack it. Payloads that - were sampled out leave hashes without a stored record. + were sampled out leave hashes without a stored record. The current hash should be put on other telemetry items, even if the sdk_config has not been sent yet. - It is understood that this means that the final hash MAY differ from the one attached to early records. + It is understood that this means that the final hash MAY differ from the one attached to early records. Those may be linked by fallback key only (see below). - **Fallback key:** `release` alone is insufficient, because configuration can differ by environment and build. The composite key is bounded, human-readable, and cheap, but `environment` defaults to From 4fd3f8c0039bcb8418b34fd689588e954cf07f87 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 10:34:41 +0200 Subject: [PATCH 39/51] small tweak --- text/0162-capture-sdk-options.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 55020221..806f4be0 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -127,7 +127,7 @@ attribute in this example except `sentry.sdk_config.normalized.*`, which Relay a } ``` -The key fields (see [Appendix A](#appendix-a-payload-details) for all fields and exact rules): +The key fields (see [Appendix A](#appendix-a-payload-details) for more details on serialization and normalization): - **`sentry.sdk_config.option.`:** the **effective** configuration that the SDK runs with (after defaults, environment variables, and derived values), under native option names. Values are reduced From 2ce4c8b85a14750a2f25ec5343b36502016c2927 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 10:40:42 +0200 Subject: [PATCH 40/51] clarify --- text/0162-capture-sdk-options.md | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 806f4be0..66916960 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -129,6 +129,10 @@ attribute in this example except `sentry.sdk_config.normalized.*`, which Relay a The key fields (see [Appendix A](#appendix-a-payload-details) for more details on serialization and normalization): +- **`sentry.sdk.packages`:** the Sentry packages that make up the SDK, the same data as `sdk.packages` + on events today. Each entry is `@`, where `name` is prefixed with its package + registry as on events (`npm:@sentry/node@10.0.0`, `pypi:sentry-sdk@2.0.0`); consumers split on the + last `@`. - **`sentry.sdk_config.option.`:** the **effective** configuration that the SDK runs with (after defaults, environment variables, and derived values), under native option names. Values are reduced to attribute types: nested objects become dot-notation keys, callbacks become `"[Function]"`, and @@ -138,7 +142,7 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for more details o `init()`, to tell actual usage apart from defaults. This should be best-effort - it MAY be incomplete if users add configuration in alternate paths or similar. If it is not possible to enumerate options automatically, SDKs MAY send a hand-picked subset of options here only. -- **Integrations:** `sentry.sdk.integrations` lists every registered integration. Per integration, +- **`sentry.sdk.integrations`:** lists every registered integration. Per integration, `sentry.sdk_config.integration..option.` holds its serialized options, and an optional `sentry.sdk_config.integration..applied` records whether it took effect at runtime. For example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically From 2308298a4a85ff434c65f228113788fb4eb6d4ae Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 13:05:03 +0200 Subject: [PATCH 41/51] small tweak --- text/0162-capture-sdk-options.md | 15 ++++++++------- 1 file changed, 8 insertions(+), 7 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 66916960..683b295c 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -170,13 +170,14 @@ New attributes to add to Sentry conventions: Reused as-is: `sentry.sdk.name`, `sentry.sdk.version`, `sentry.sdk.integrations`, `sentry.release`, `sentry.environment`, `sentry.dist`, `process.runtime.name`, `process.runtime.version`. -The design keeps SDKs simple: they serialize their existing options object (e.g. JS -`client.getOptions()`), diff its keys against the `init()` argument, and ship no name mapping. New options are captured -automatically, one Relay implementation avoids inconsistent mappings across SDKs, and the catalog can -change, including for stored data, without SDK releases. Sending two option trees or wrapping every -value as `{ value, source }` would need more SDK bookkeeping. The cost is that converted values lose -their original form: for example, we cannot tell whether `integrations` was passed as a function (in -JS, only `integrations` and `stackParser` are affected). +The design keeps SDKs simple: most SDKs can just serialize their existing finalized options object (e.g. JS +`client.getOptions()`), new options are captured automatically. +One Relay implementation avoids inconsistent mappings across SDKs, and the catalog can +change, including for stored data, without SDK releases. + +One downside of this is that converted values lose their original form: +for example, we cannot tell whether `integrations` was passed as a function - +this should generally only affect a small subset of options (e.g. in JS, only `integrations` and `stackParser` are affected). Sensitive data is scrubbed primarily server-side. SDKs MAY also scrub values that they know to be sensitive, but they do not guarantee fully scrubbed data. From aafa26df7847d1e00b89f24d0c720eee8d6654f3 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 13:08:14 +0200 Subject: [PATCH 42/51] small clarification --- text/0162-capture-sdk-options.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 683b295c..21e6239d 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -211,7 +211,7 @@ release separate. If no stored record matches an event's hash (e.g. because the payload was lost or sampled out), or the event has no hash, correlation falls back to `release` + `environment` + `dist`, which every event already carries. The fallback resolves only to the records of that combination and is coarse -when `release` is unset. +when `release` is unset. It should use the newest matching config in this case. # Drawbacks From 7b3adf12bb899ea2965f16872277133f494218f9 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 13:17:02 +0200 Subject: [PATCH 43/51] small ref --- text/0162-capture-sdk-options.md | 1 + 1 file changed, 1 insertion(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 21e6239d..75d06d6a 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -238,6 +238,7 @@ when `release` is unset. It should use the newest matching config in this case. expire after their retention period, so long-running processes must re-send within it. EAP could deduplicate via a deterministic `item_id` from the hash plus a bucketed `timestamp` (as preprod does), to be confirmed with the EAP team. Limits on attribute count and size per item. +- **Custom filtering:** do we need SDK side filtering, e.g. `beforeSendSdkConfig`, to allow manual PII stripping? ## Out of scope From 4e297dd1d4d84d0ac210431920b494a041b0790e Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Wed, 30 Sep 2026 13:29:05 +0200 Subject: [PATCH 44/51] hash instructions --- text/0162-capture-sdk-options.md | 3 ++- 1 file changed, 2 insertions(+), 1 deletion(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 75d06d6a..8058ca14 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -206,7 +206,8 @@ SDKs MUST compute the hash from the serialized option and integration attributes to every event: as the same `sentry.sdk_config.hash` attribute on spans, logs, and other items with attributes, and in a new `sdk_config.hash` context field on errors and transactions. The hash links each event to its exact configuration and keeps configurations that differ within one -release separate. +release separate. The hash MAY be generated based off the full `attributes` hash (minus the `sentry.sdk_config.hash` attribute), +or from a subset if that makes more sense for an SDK. If no stored record matches an event's hash (e.g. because the payload was lost or sampled out), or the event has no hash, correlation falls back to `release` + `environment` + `dist`, which every From 93a0bcdefef4cf166dd110549fe8b191e2fe22a9 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 08:22:30 +0200 Subject: [PATCH 45/51] updates --- text/0162-capture-sdk-options.md | 30 ++++++++++++++++++------------ 1 file changed, 18 insertions(+), 12 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 8058ca14..d9c4be3d 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -87,6 +87,10 @@ attribute in this example except `sentry.sdk_config.normalized.*`, which Relay a "type": "array", "value": ["InboundFilters", "Express", "Fastify", "Koa", "MyIntegration"] }, + "sentry.sdk.integrations.applied": { + "type": "array", + "value": ["Express"] + }, "sentry.release": { "type": "string", "value": "my-app@1.2.3" }, "sentry.environment": { "type": "string", "value": "production" }, "sentry.dist": { "type": "string", "value": "42" }, @@ -108,8 +112,6 @@ attribute in this example except `sentry.sdk_config.normalized.*`, which Relay a "value": ["dsn", "tracesSampleRate", "sendDefaultPii", "beforeSend", "denyUrls"] }, - "sentry.sdk_config.integration.Express.applied": { "type": "boolean", "value": true }, - "sentry.sdk_config.integration.Fastify.applied": { "type": "boolean", "value": false }, "sentry.sdk_config.integration.MyIntegration.option.filter": { "type": "string", "value": "aaa" }, "sentry.sdk_config.integration.MyIntegration.option.shouldLog": { "type": "string", "value": "[Function]" }, @@ -143,10 +145,10 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for more details o incomplete if users add configuration in alternate paths or similar. If it is not possible to enumerate options automatically, SDKs MAY send a hand-picked subset of options here only. - **`sentry.sdk.integrations`:** lists every registered integration. Per integration, - `sentry.sdk_config.integration..option.` holds its serialized options, and an optional - `sentry.sdk_config.integration..applied` records whether it took effect at runtime. For - example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically - uses only one. + `sentry.sdk_config.integration..option.` holds its serialized options (in dot-nested notation). +- **`sentry.sdk.integrations.applied`:** This optional array attribute records all integrations that we + specifically want to track for them having been applied at runtime. + For example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. - **`sentry.sdk_config.hash`:** a required hash of the configuration (see [Storage](#storage)). - **`sentry.sdk_config.normalized.`:** a small catalog of options under canonical cross-SDK names (JS `tracesSampleRate` → `traces_sample_rate`), derived by Relay. @@ -158,11 +160,11 @@ New attributes to add to Sentry conventions: | Attribute | Type | Example | | --------------------------------------------------- | -------- | ---------------------------------------------------------------- | | `sentry.sdk.packages` | string[] | `["npm:@sentry/node@10.0.0"]` | +| `sentry.sdk.integrations.applied` | string[] | `["Express"]` | | `sentry.sdk_config.version` | integer | `1` | | `sentry.sdk_config.hash` | string | `"9f2c1a7e"` | | `sentry.sdk_config.option.` | any | `sentry.sdk_config.option.sampleRate=1.0` | | `sentry.sdk_config.options_set_by_user` | string[] | `["dsn", "tracesSampleRate"]` | -| `sentry.sdk_config.integration..applied` | boolean | `true` | | `sentry.sdk_config.integration..option.` | any | `...MyIntegration.option.filter="aaa"` | | `sentry.sdk_config.normalized.` | any | `...normalized.traces_sample_rate=0.2` | | `sentry.sdk_config.normalized..original` | string | `sentry.sdk_config.normalized.sample_rate.original="sampleRate"` | @@ -185,7 +187,7 @@ sensitive, but they do not guarantee fully scrubbed data. ## Sending We propose to send one payload per `init()` once the configuration has -settled, because values such as `applied` are only known after `init()` returns. Each SDK MAY choose +settled, because values such as `sentry.sdk.integrations.applied` are only known after `init()` returns. Each SDK MAY choose how to wait, for example with a short debounce (e.g. 5 seconds) or a lifecycle hook. Short-lived processes MAY flush the payload on shutdown. [Appendix B](#appendix-b-sending-and-storage-details) has the details. @@ -206,7 +208,7 @@ SDKs MUST compute the hash from the serialized option and integration attributes to every event: as the same `sentry.sdk_config.hash` attribute on spans, logs, and other items with attributes, and in a new `sdk_config.hash` context field on errors and transactions. The hash links each event to its exact configuration and keeps configurations that differ within one -release separate. The hash MAY be generated based off the full `attributes` hash (minus the `sentry.sdk_config.hash` attribute), +release separate. The hash MAY be generated based off the full `attributes` object (minus the `sentry.sdk_config.hash` attribute), or from a subset if that makes more sense for an SDK. If no stored record matches an event's hash (e.g. because the payload was lost or sampled out), or @@ -214,6 +216,10 @@ the event has no hash, correlation falls back to `release` + `environment` + `di event already carries. The fallback resolves only to the records of that combination and is coarse when `release` is unset. It should use the newest matching config in this case. +The new data type should be free to end users, we do not plan on billing for it. +The estimated amount of envelopes to be sent equals the amount of session envelopes we get today, +as generally a session is sent per application startup. + # Drawbacks - **Sensitive data:** configuration can contain secrets and PII (DSNs, tokens or headers in transport @@ -224,7 +230,7 @@ when `release` is unset. It should use the newest matching config in this case. picture, for example when deciding whether an option "looks unused". - **Lossy callbacks:** `"[Function]"` shows that filtering is configured, not what it filters. - **Maintenance:** the normalization catalog must track options across all SDKs. -- **SDK complexity:** delayed sending, flushing, `applied` tracking, serialization, and sampling add +- **SDK complexity:** delayed sending, flushing, `sentry.sdk.integrations.applied` tracking, serialization, and sampling add code and runtime cost on every `init()`. - **Deduplication identity:** a poor hash or a missing `release` either stores too much or merges distinct configurations. @@ -287,7 +293,7 @@ deterministic and consistent across SDKs: # Appendix B: Sending and storage details -- **Waiting for settled configuration:** values known only after `init()` include `applied`, lazily +- **Waiting for settled configuration:** values known only after `init()` include `sentry.sdk.integrations.applied`, lazily registered integrations, and an asynchronously detected release or environment. The debounce MAY be configurable for setups that settle later. A lifecycle hook that guarantees settled options (e.g. after the first request, a post-init hook, or the first idle event loop) is preferable. The goal is @@ -299,7 +305,7 @@ deterministic and consistent across SDKs: persisted. To accomodate this, as well as retention period dropping config after longer time periods, a single lost send would leave a long-running server without stored configuration. - **Hash:** identical configurations produce identical hashes, and any change to the options, the - registered integrations, their options, or `applied` changes the hash. Algorithms do not need to + registered integrations, their options, or the applied integrations changes the hash. Algorithms do not need to match across SDKs. The hash is computed SDK-side so that events can carry it. - **Hash caveats:** it adds a field to every event. Because configuration settles after `init()`, the hash on events must match the reported configuration (e.g. by hashing only values that are stable From 17a54bab0813cad1e96a2407ece684ff8fe15e3b Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 09:12:47 +0200 Subject: [PATCH 46/51] prettier and add estimations --- text/0162-capture-sdk-options.md | 54 ++++++++++++++++++++++++++------ 1 file changed, 45 insertions(+), 9 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index d9c4be3d..b02c515c 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -146,8 +146,8 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for more details o enumerate options automatically, SDKs MAY send a hand-picked subset of options here only. - **`sentry.sdk.integrations`:** lists every registered integration. Per integration, `sentry.sdk_config.integration..option.` holds its serialized options (in dot-nested notation). -- **`sentry.sdk.integrations.applied`:** This optional array attribute records all integrations that we - specifically want to track for them having been applied at runtime. +- **`sentry.sdk.integrations.applied`:** This optional array attribute records all integrations that we + specifically want to track for them having been applied at runtime. For example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. - **`sentry.sdk_config.hash`:** a required hash of the configuration (see [Storage](#storage)). - **`sentry.sdk_config.normalized.`:** a small catalog of options under canonical cross-SDK @@ -175,10 +175,10 @@ Reused as-is: `sentry.sdk.name`, `sentry.sdk.version`, `sentry.sdk.integrations` The design keeps SDKs simple: most SDKs can just serialize their existing finalized options object (e.g. JS `client.getOptions()`), new options are captured automatically. One Relay implementation avoids inconsistent mappings across SDKs, and the catalog can -change, including for stored data, without SDK releases. +change, including for stored data, without SDK releases. -One downside of this is that converted values lose their original form: -for example, we cannot tell whether `integrations` was passed as a function - +One downside of this is that converted values lose their original form: +for example, we cannot tell whether `integrations` was passed as a function - this should generally only affect a small subset of options (e.g. in JS, only `integrations` and `stackParser` are affected). Sensitive data is scrubbed primarily server-side. SDKs MAY also scrub values that they know to be @@ -208,7 +208,7 @@ SDKs MUST compute the hash from the serialized option and integration attributes to every event: as the same `sentry.sdk_config.hash` attribute on spans, logs, and other items with attributes, and in a new `sdk_config.hash` context field on errors and transactions. The hash links each event to its exact configuration and keeps configurations that differ within one -release separate. The hash MAY be generated based off the full `attributes` object (minus the `sentry.sdk_config.hash` attribute), +release separate. The hash MAY be generated based off the full `attributes` object (minus the `sentry.sdk_config.hash` attribute), or from a subset if that makes more sense for an SDK. If no stored record matches an event's hash (e.g. because the payload was lost or sampled out), or @@ -216,9 +216,8 @@ the event has no hash, correlation falls back to `release` + `environment` + `di event already carries. The fallback resolves only to the records of that combination and is coarse when `release` is unset. It should use the newest matching config in this case. -The new data type should be free to end users, we do not plan on billing for it. -The estimated amount of envelopes to be sent equals the amount of session envelopes we get today, -as generally a session is sent per application startup. +The new data type should be free to end users, we do not plan on billing for it. +See [Appendix C](#appendix-c-volume-estimates) for estimations. # Drawbacks @@ -318,3 +317,40 @@ deterministic and consistent across SDKs: and build. The composite key is bounded, human-readable, and cheap, but `environment` defaults to `production` and `dist` is usually absent, so the key depends on `release`, which is often unset. It resolves to the first-seen record or to the set of stored variants. + +# Appendix C: Volume Estimates + +We do not neatly track anything that proxies to number of `init()` calls, which would be the rough equivalent +of volume we expect for this feature. We can approximate this a bit by looking at browser & mobile projects, +where a session equals to an `init()` call, generally. The following data is for active projects in a single day: + +| Platform | Active projects | # Sessions | # Sessions / project | +| -------- | --------------- | ---------- | -------------------- | +| Browser | 127k | 29b | 239k | +| Mobile | 80k | 5b | 64k | + +For the sake of estimation, we can scale both the mobile and browser averages up for server SDK active projects: + +| Platform | Active projects | Est. lower bound # sessions | Est. upper bound # sessions | +| ---------- | --------------- | --------------------------- | --------------------------- | +| Server | 82k | 5b | 19b | +| Desktop | 2k | 153m | 574m | +| Serverless | 1k | 72m | 271m | + +**Total Estimated init calls per day:** based on this, the estimate would be 40-54 billion/day. + +NOTE: This is likely a very high estimate, because many/most server projects will have considerably less init calls/release than client SDKs. + +Another data point to be used: # of releases per day: + +- Browser: ~886K +- Mobile: ~825K +- Unmapped SDK: ~406K +- Server: ~204K +- Desktop: ~38K +- Browser+mobile (hybrid): ~27K +- Serverless: ~2.7K + +For server SDKs, this may be closer to the number of init calls then the session estimation above. + +Combining these two datasets, a reasonable estimation for **init cals per day** could be _~35b_ From 5d330a41765a82514a4511f168190e145b53dce1 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 09:18:52 +0200 Subject: [PATCH 47/51] instructions for convention attributes --- text/0162-capture-sdk-options.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index b02c515c..7ab793b7 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -169,6 +169,8 @@ New attributes to add to Sentry conventions: | `sentry.sdk_config.normalized.` | any | `...normalized.traces_sample_rate=0.2` | | `sentry.sdk_config.normalized..original` | string | `sentry.sdk_config.normalized.sample_rate.original="sampleRate"` | +These should be added as internal attributes, with a note that they are only supposed to be used in the sdk_config item type. + Reused as-is: `sentry.sdk.name`, `sentry.sdk.version`, `sentry.sdk.integrations`, `sentry.release`, `sentry.environment`, `sentry.dist`, `process.runtime.name`, `process.runtime.version`. From a180365e8352e4d5e23f582dab4d0a1c69dcdc75 Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 09:51:41 +0200 Subject: [PATCH 48/51] better estimates --- text/0162-capture-sdk-options.md | 38 +++++--------------------------- 1 file changed, 5 insertions(+), 33 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 7ab793b7..268b61ab 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -322,37 +322,9 @@ deterministic and consistent across SDKs: # Appendix C: Volume Estimates -We do not neatly track anything that proxies to number of `init()` calls, which would be the rough equivalent -of volume we expect for this feature. We can approximate this a bit by looking at browser & mobile projects, -where a session equals to an `init()` call, generally. The following data is for active projects in a single day: +We do not neatly track the number of `init()` calls, which would be the rough equivalent +of volume we expect for this feature. We can approximate this by looking at session & session aggregates being sent. +Based on this, a naive expectation would be about **50b** events being sent per day (before deduplication). -| Platform | Active projects | # Sessions | # Sessions / project | -| -------- | --------------- | ---------- | -------------------- | -| Browser | 127k | 29b | 239k | -| Mobile | 80k | 5b | 64k | - -For the sake of estimation, we can scale both the mobile and browser averages up for server SDK active projects: - -| Platform | Active projects | Est. lower bound # sessions | Est. upper bound # sessions | -| ---------- | --------------- | --------------------------- | --------------------------- | -| Server | 82k | 5b | 19b | -| Desktop | 2k | 153m | 574m | -| Serverless | 1k | 72m | 271m | - -**Total Estimated init calls per day:** based on this, the estimate would be 40-54 billion/day. - -NOTE: This is likely a very high estimate, because many/most server projects will have considerably less init calls/release than client SDKs. - -Another data point to be used: # of releases per day: - -- Browser: ~886K -- Mobile: ~825K -- Unmapped SDK: ~406K -- Server: ~204K -- Desktop: ~38K -- Browser+mobile (hybrid): ~27K -- Serverless: ~2.7K - -For server SDKs, this may be closer to the number of init calls then the session estimation above. - -Combining these two datasets, a reasonable estimation for **init cals per day** could be _~35b_ +In regard to storage, there are about **2.5m** releases per day, which could roughly translate to the number of +_stored config payloads_. From 4d48b266976f69b6b055f7e8173fbaea5351fefc Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 10:08:14 +0200 Subject: [PATCH 49/51] adjust options --- text/0162-capture-sdk-options.md | 48 +++++++++++++++++--------------- 1 file changed, 25 insertions(+), 23 deletions(-) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 268b61ab..248b7f05 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -83,14 +83,7 @@ attribute in this example except `sentry.sdk_config.normalized.*`, which Relay a "sentry.sdk.name": { "type": "string", "value": "sentry.javascript.node" }, "sentry.sdk.version": { "type": "string", "value": "10.0.0" }, "sentry.sdk.packages": { "type": "array", "value": ["npm:@sentry/node@10.0.0"] }, - "sentry.sdk.integrations": { - "type": "array", - "value": ["InboundFilters", "Express", "Fastify", "Koa", "MyIntegration"] - }, - "sentry.sdk.integrations.applied": { - "type": "array", - "value": ["Express"] - }, + "sentry.release": { "type": "string", "value": "my-app@1.2.3" }, "sentry.environment": { "type": "string", "value": "production" }, "sentry.dist": { "type": "string", "value": "42" }, @@ -112,8 +105,16 @@ attribute in this example except `sentry.sdk_config.normalized.*`, which Relay a "value": ["dsn", "tracesSampleRate", "sendDefaultPii", "beforeSend", "denyUrls"] }, - "sentry.sdk_config.integration.MyIntegration.option.filter": { "type": "string", "value": "aaa" }, - "sentry.sdk_config.integration.MyIntegration.option.shouldLog": { "type": "string", "value": "[Function]" }, + "sentry.sdk.integrations": { + "type": "array", + "value": ["InboundFilters", "Express", "Fastify", "Koa", "MyIntegration"] + }, + "sentry.sdk.integrations.applied": { + "type": "array", + "value": ["Express"] + }, + "sentry.sdk.integrations.options.MyIntegration.filter": { "type": "string", "value": "aaa" }, + "sentry.sdk.integrations.options.MyIntegration.shouldLog": { "type": "string", "value": "[Function]" }, "sentry.sdk_config.normalized.sample_rate": { "type": "double", "value": 1.0 }, "sentry.sdk_config.normalized.sample_rate.original": { "type": "string", "value": "sampleRate" }, @@ -144,11 +145,12 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for more details o `init()`, to tell actual usage apart from defaults. This should be best-effort - it MAY be incomplete if users add configuration in alternate paths or similar. If it is not possible to enumerate options automatically, SDKs MAY send a hand-picked subset of options here only. -- **`sentry.sdk.integrations`:** lists every registered integration. Per integration, - `sentry.sdk_config.integration..option.` holds its serialized options (in dot-nested notation). +- **`sentry.sdk.integrations`:** lists every registered integration. - **`sentry.sdk.integrations.applied`:** This optional array attribute records all integrations that we specifically want to track for them having been applied at runtime. For example, the Node SDK registers Express, Fastify, Koa, and more by default, but an app typically uses only one. +- **`sentry.sdk.integrations.options..`:** Per integration, holds its serialized options (in dot-nested notation). + This MAY be captured for some or all integration options. SDKs MAY decide which integration options are relevant to capture. - **`sentry.sdk_config.hash`:** a required hash of the configuration (see [Storage](#storage)). - **`sentry.sdk_config.normalized.`:** a small catalog of options under canonical cross-SDK names (JS `tracesSampleRate` → `traces_sample_rate`), derived by Relay. @@ -157,17 +159,17 @@ The key fields (see [Appendix A](#appendix-a-payload-details) for more details o New attributes to add to Sentry conventions: -| Attribute | Type | Example | -| --------------------------------------------------- | -------- | ---------------------------------------------------------------- | -| `sentry.sdk.packages` | string[] | `["npm:@sentry/node@10.0.0"]` | -| `sentry.sdk.integrations.applied` | string[] | `["Express"]` | -| `sentry.sdk_config.version` | integer | `1` | -| `sentry.sdk_config.hash` | string | `"9f2c1a7e"` | -| `sentry.sdk_config.option.` | any | `sentry.sdk_config.option.sampleRate=1.0` | -| `sentry.sdk_config.options_set_by_user` | string[] | `["dsn", "tracesSampleRate"]` | -| `sentry.sdk_config.integration..option.` | any | `...MyIntegration.option.filter="aaa"` | -| `sentry.sdk_config.normalized.` | any | `...normalized.traces_sample_rate=0.2` | -| `sentry.sdk_config.normalized..original` | string | `sentry.sdk_config.normalized.sample_rate.original="sampleRate"` | +| Attribute | Type | Example | +| ---------------------------------------------- | -------- | ---------------------------------------------------------------- | +| `sentry.sdk.packages` | string[] | `["npm:@sentry/node@10.0.0"]` | +| `sentry.sdk.integrations.applied` | string[] | `["Express"]` | +| `sentry.sdk.integrations.options..` | any | `...integrations.options.MyIntegration.filter="aaa"` | +| `sentry.sdk_config.version` | integer | `1` | +| `sentry.sdk_config.hash` | string | `"9f2c1a7e"` | +| `sentry.sdk_config.option.` | any | `sentry.sdk_config.option.sampleRate=1.0` | +| `sentry.sdk_config.options_set_by_user` | string[] | `["dsn", "tracesSampleRate"]` | +| `sentry.sdk_config.normalized.` | any | `...normalized.traces_sample_rate=0.2` | +| `sentry.sdk_config.normalized..original` | string | `sentry.sdk_config.normalized.sample_rate.original="sampleRate"` | These should be added as internal attributes, with a note that they are only supposed to be used in the sdk_config item type. From 1ff61af369e2a4a7065e845a34877299ba57a40e Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 10:29:28 +0200 Subject: [PATCH 50/51] note envelope header --- text/0162-capture-sdk-options.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 248b7f05..04a5ce92 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -215,6 +215,8 @@ hash links each event to its exact configuration and keeps configurations that d release separate. The hash MAY be generated based off the full `attributes` object (minus the `sentry.sdk_config.hash` attribute), or from a subset if that makes more sense for an SDK. +The hash also MUST be part of the envelope item header and we MUST reject any item without a hash. + If no stored record matches an event's hash (e.g. because the payload was lost or sampled out), or the event has no hash, correlation falls back to `release` + `environment` + `dist`, which every event already carries. The fallback resolves only to the records of that combination and is coarse From e2be71a8821ad8b64a4329c3015ae2242fd8499b Mon Sep 17 00:00:00 2001 From: Francesco Novy Date: Thu, 1 Oct 2026 12:49:44 +0200 Subject: [PATCH 51/51] note for options --- text/0162-capture-sdk-options.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/text/0162-capture-sdk-options.md b/text/0162-capture-sdk-options.md index 04a5ce92..f4263d83 100644 --- a/text/0162-capture-sdk-options.md +++ b/text/0162-capture-sdk-options.md @@ -27,6 +27,8 @@ With it, we can support: settings. - **Discoverability of data sources:** show every `Sentry.init()` that sends data into a project, with its configuration over time, and eventually allow changing configuration from the UI. +- **Onboarding**: We could leverage this information to make onboarding easier - we could then know with certainty + if the SDK setup was successfull, no need to ask users to send logs/metrics/spans/errors to verify SDK setup. # Background