Datadog Event Forwarding
Ship the app's metrics, logs, and deploy events to Datadog with tags that stay consistent.
What it adds
A single outbound telemetry path to Datadog, with an agreed tag set, batching, sampling, and redaction applied before send.
What your agent is told to do
5
What your agent is told to do
5-
1
Decide first what is worth sending — the handful of metrics, log streams, and deployment markers someone would actually look at during an incident — and route all of them through one outbound path rather than scattering calls across the codebase.
-
2
Fix a small, mandatory tag set applied to everything: service, environment, version, and where the app is multi-tenant, an opaque tenant identifier. Set them from the same configuration the deploy uses, so a dashboard filter works the same across metrics, logs, and traces.
-
3
Treat every tag value as a cardinality decision. User identifiers, request identifiers, URLs with embedded values, and raw error messages must not become tag values; put them in the log body or trace attributes where they belong.
-
4
Buffer and batch outbound telemetry through the app's existing background processing, with a bounded queue that drops oldest-first when it fills. Telemetry must never be sent inline on the request path or hold a user request open.
-
5
Redact before the payload is constructed, not after: credentials, tokens, authorisation headers, payment details, and personal data must never reach the buffer. Keep the API credential in server-side configuration and never in a browser bundle.
Edge cases it handles
8
Edge cases it handles
8- High-cardinality tags are the standard way this integration becomes expensive overnight, because each distinct value creates a new time series. Enforce an allowed tag key list and reject values that look unbounded.
- Payloads also have size limits. Truncate long log bodies and stack traces at a defined length with a marker, rather than letting an oversized payload be rejected wholesale and the event lost.
- Inconsistent service, environment, version, or tenant metadata makes correlation impossible — a metric tagged one way and a log tagged another cannot be joined. Derive all four from one place and assert their presence before send.
- Secrets and personal data leak in through log lines and exception messages more often than through deliberate fields. Redact at the source and default to excluding unknown structured fields.
- A noisy loop or a failing endpoint will flood the pipeline and the bill. Sample repetitive events at a defined rate and aggregate counters locally before shipping them rather than sending one call per occurrence.
- When Datadog is unreachable, the app must degrade to local logging and continue serving users. Do not retry synchronously, do not grow the buffer without limit, and do not let a telemetry failure surface as a user-facing error.
- Backoff must be applied on rate limiting and on server errors, with jitter, or every instance will retry in lockstep and prolong the outage.
- Losing telemetry is acceptable; losing user requests is not. Make the dropped-event count itself an observable number so silent loss is visible.
Definition of done
8
Definition of done
8- All outbound telemetry leaves through a single path with a mandatory service, environment, version, and tenant tag set.
- Tag keys are restricted to an allowed list and unbounded values are rejected before send.
- No credential or personal data is present in forwarded metrics, logs, or events, verified against a captured payload.
- Sending is batched off the request path and never delays or fails a user request.
- Noisy events are sampled or aggregated, and oversized payloads are truncated rather than dropped whole.
- A provider outage falls back to local logging with a bounded buffer, backoff with jitter, and a visible dropped-event count.
- The feature matches the existing design system.
- No existing functionality is broken.
Related features
Multi-Model Routing
Multi-Model Routing
Send each AI request to the right model using rules you can read and test.
What it does
A deterministic routing layer that picks a model per request from task type, context size, latency budget, and data sensitivity.
How it works
- 1 Express routing as explicit, ordered rules over inputs the app can measure: task type, estimated context size, latency budget, and the sensitivity classification of the data involved. A rule set that can be read line by line can be reviewed and tested.
- 2 Make routing deterministic. The same inputs must always produce the same route, so a bad output can be reproduced and a rule change can be evaluated. Randomised or load-based selection turns every incident into guesswork.
- 3 Classify data before routing and refuse to route restricted content to any destination not approved for it. This check is a hard block, not a preference, and it must run before the request is assembled.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/multi-model-routing
AI Cost Budgets
AI Cost Budgets
Cap what AI features are allowed to spend before the bill arrives.
What it does
Monetary spending limits on AI work, scoped by workspace, feature, and time period, enforced before a run starts.
How it works
- 1 Find every place the app calls a model and route all of them through one accounting point that records estimated and actual spend against a scope. A budget that only covers the chat feature is not a budget.
- 2 Estimate the cost of a run from the size of its input before dispatching it, and refuse anything that would exceed the remaining budget on its own.
- 3 Reserve the estimate against the budget when the run starts, then reconcile to the real usage figures when it finishes, releasing whatever was over-reserved.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/ai-cost-budgets
Model Selection
Model Selection
Let each AI task run on the model that suits its quality, speed, and cost needs.
What it does
A per-task model choice, drawn from the models the app already has configured, with capability filtering and safe defaults.
How it works
- 1 Enumerate the models the app already has access to and record what each one can actually do: context capacity, whether it can return the structured output the task requires, whether it supports the tools the task calls, and its relative cost and speed.
- 2 Offer only the models that satisfy the task's requirements. A task that needs structured output must not list a model that cannot reliably produce it, because the failure appears later as malformed responses rather than as an unavailable option.
- 3 Store the choice against the specific task, not as one global setting. A single default forces a summarisation task and a classification task onto the same tier when they have opposite needs.
Copy the prompt
No account needed
Add this feature to my app:
https://addthisfeature.com/x/model-selection
How it works
-
1
Copy the link
Grab the Markdown instruction URL for this feature.
-
2
Give it to your AI
Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.
-
3
It inspects, then implements
Your agent reads your existing app first, then adds the feature to fit it.
Works with your stack
These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.
Need it tighter than that? Customize the feature and tell it exactly what you're running.