How to Keep Secrets and Sensitive Data Out of AI Prompts, Embeddings, Traces, and Logs

TL;DR

Sensitive data protection for AI is not a prompt-writing problem. It is a data-path control problem.

The safest operating model classifies data before it enters the AI workflow, blocks credentials and restricted records by default, redacts or tokenizes approved confidential data, preserves source permissions in retrieval, disables prompt and tool-content telemetry unless explicitly approved, scans repositories and prompt assets for secrets, encrypts every persistent store, and gives each data surface a defined retention and deletion path.

When a leak occurs, treat the exposed value as compromised. Stop capture, revoke or rotate credentials, quarantine affected indexes and telemetry, determine every replicated copy, purge or rebuild the contaminated data set, and validate that the system cannot reproduce the exposure before service is restored.

Introduction

An AI application can leak sensitive data without ever suffering a traditional database breach.

A developer pastes a production token into a troubleshooting prompt. A retrieval pipeline embeds a restricted document into a shared vector index. An agent writes tool arguments into a trace. A debug logger records the complete prompt and response. A support engineer exports telemetry to a ticket. A backup preserves data that the primary system later deletes.

Each individual action may look operationally convenient. Together, they create a distributed sensitive-data problem across systems that were rarely designed under one retention, access, and incident-response model.

This is why “do not paste secrets into AI” is necessary but insufficient. User training cannot protect a workflow that automatically copies raw content into prompts, embeddings, caches, traces, logs, evaluation sets, and long-lived observability platforms.

A production design must assume that sensitive data will eventually approach the AI boundary. The architecture should decide what happens before the model, vector database, or telemetry backend receives it.

Scenario and Operational Risk

Assume an enterprise AI assistant can answer operational questions using internal documentation and call approved tools. The application uses retrieval-augmented generation, distributed tracing, centralized logs, a prompt repository, and an evaluation pipeline.

During an incident review, the security team discovers that a complete database connection string appeared in a trace attribute. The same value was also present in the original user prompt, an application log, a failed-request dead-letter queue, and a copied troubleshooting ticket.

The first discovery is not the full incident. It is only the first confirmed location.

AI data exposure expands quickly because one request can create multiple replicas:

Surface
How sensitive data arrives
Why it persists

Prompt and conversation state
User input, system instructions, memory, retrieved context
Session history, caches, provider retention, debugging

Embeddings and vector stores
Documents, tickets, source code, records, chat transcripts
Indexes, replicas, snapshots, stale chunks

Tool calls
Arguments, results, API errors, file content
Traces, audit records, retry queues

Traces and logs
Automatic instrumentation, exception capture, request bodies
Central retention, exports, dashboards, archives

Evaluation data
Failed examples, human review sets, regression cases
Versioned test corpora and experiment tracking

Build and source systems
Prompt templates, notebooks, configuration, test fixtures
Git history, pull requests, issue comments, artifacts

Backups and exports
Snapshots, support bundles, analyst downloads
Independent retention and delayed deletion

The operational risk is not limited to disclosure. Sensitive data can also be retrieved by the wrong user, retained beyond policy, copied into a lower-trust environment, exposed to a third-party processor, or preserved after the source record was deleted.

Protection Scope and Assumptions

This runbook applies to production and preproduction AI systems that use one or more of the following:

hosted or self-managed language models

retrieval-augmented generation

vector databases or embedding services

AI agents and tool calls

prompt, workflow, or evaluation repositories

centralized logs, traces, and metrics

data loss prevention or redaction services

CI/CD pipelines and secret scanning

The runbook assumes the organization already has an identity provider, a secrets manager, encryption key management, incident response, data ownership, and a basic information-classification policy. The AI platform should consume those controls rather than invent weaker parallel versions.

This is a technical operating model, not a substitute for legal, privacy, records-management, or regulatory guidance. Retention periods, breach-notification requirements, and approved processing locations must be set by the appropriate organizational owners.

The Sensitive Data Control Path

The most important architectural decision is where policy enforcement occurs. A warning in the user interface is advisory. A system prompt telling the model not to expose data is probabilistic. A gateway that blocks or transforms content before it reaches the model is enforceable.

The reader should notice that classification and policy decisions occur before prompt assembly, embedding generation, and telemetry export. The model is not the security boundary.

A mature implementation repeats inspection at more than one point. Input controls stop sensitive content before processing. Retrieval controls prevent unauthorized context from being selected. Output controls catch generated or retrieved disclosures. Telemetry controls prevent diagnostic systems from becoming the final leak destination.

Build a Classification Model That AI Systems Can Enforce

A policy that says “protect sensitive data” is not executable. The AI data path needs a small, testable classification model with a defined action for each class.

A practical baseline is:

Classification
Typical examples
Default AI action

Public
Published documentation, approved public content
Allow

Internal
Routine procedures, non-sensitive architecture notes
Allow in approved environments with normal access controls

Confidential
Customer records, internal financial data, employee information, proprietary code
Redact, tokenize, minimize, or require an approved use case

Restricted
Regulated records, high-impact personal data, legal hold material, sensitive investigations
Deny unless a specifically approved architecture exists

Credential
Passwords, private keys, access tokens, connection strings, signing material
Deny, alert, revoke or rotate if exposure is real

Do not classify only by document. Classification can change at the field, chunk, query, and output level. A public runbook can still contain a copied credential. An internal ticket can contain a restricted customer record. A tool result can combine several classes in one response.

Assign an Action, Owner, and Evidence Requirement

Each classification needs three operational properties:

Action: allow, redact, tokenize, block, quarantine, or require approval

Owner: application, data owner, platform, security, privacy, or records management

Evidence: policy decision, detector result, exception ID, deletion proof, or review record

Without ownership, every team assumes another layer is filtering the data. Without evidence, the organization cannot prove that a control ran or determine why an exception was allowed.

Put DLP and Redaction Before Prompt Assembly

Prompt assembly is where system instructions, user messages, conversation memory, retrieved documents, and tool results become one model request. It is also the last reliable point to stop data before it leaves the application boundary.

Inspect each input independently before concatenation. That preserves provenance and makes policy decisions explainable.

Separate Secrets from Instructions

Never place credentials, private endpoints, signing material, database passwords, or authorization rules inside a system prompt. A system prompt is model input, not a secrets manager or policy engine.

The safer pattern is:

store credentials in a secrets manager

provide them only to the tool runtime that needs them

keep the secret out of model-visible arguments and results

enforce authorization outside the model

return only the minimum business result to the agent

An agent that calls a billing tool should receive the approved invoice status, not the tool’s bearer token, database connection string, or raw backend response.

Choose the Transformation by Use Case

Redaction is not always the right transformation.

Block credentials and data that has no approved AI use.

Mask values when partial display is needed for a human reviewer.

Tokenize identities when the workflow needs consistent correlation without the original value.

Generalize precise attributes when a broader category is sufficient.

Encrypt data only when the workflow requires controlled reversibility and key ownership is clear.

Replace with typed labels when the model needs semantic context but not the original value.

For example, replacing a real customer identifier with [CUSTOMER_ID_1] can preserve the structure of a troubleshooting conversation without exposing the source value. The mapping must remain outside model and telemetry paths, in a tightly controlled store, when re-identification is required.

Use a Deny-by-Default Policy

The following vendor-neutral YAML shows the policy shape. It is not tied to a specific gateway. Replace the data classes, exception workflow, retention periods, and owner groups with organizational values.

policy_version: 1
policy_name: ai-sensitive-data-boundary

default_action: deny

classifications:
public:
prompt_action: allow
embedding_action: allow
telemetry_content: deny

internal:
prompt_action: allow
embedding_action: allow_with_source_acl
telemetry_content: deny

confidential:
prompt_action: redact_or_tokenize
embedding_action: require_approved_dataset
telemetry_content: deny

restricted:
prompt_action: deny
embedding_action: deny
telemetry_content: deny
exception_required: true

credential:
prompt_action: deny_and_alert
embedding_action: deny_and_alert
telemetry_content: deny_and_alert
incident_severity: high

telemetry:
capture_prompt_content: false
capture_response_content: false
capture_tool_arguments: false
capture_tool_results: false
allowed_attributes:
– service.name
– deployment.environment.name
– gen_ai.agent.name
– gen_ai.operation.name
– gen_ai.provider.name
– gen_ai.request.model
– gen_ai.response.model
– gen_ai.usage.input_tokens
– gen_ai.usage.output_tokens
– error.type

exceptions:
require_owner: true
require_expiration: true
require_ticket: true
maximum_duration_hours: 24

Successful enforcement means a blocked request never reaches the model, embedding service, retry queue, or content-bearing telemetry field. A policy that blocks the provider request but logs the rejected raw prompt has still failed.

Use Layered Detection Instead of One Regex

Sensitive-data detection should combine several methods because no single detector covers every data type.

A practical pipeline can include:

provider and credential patterns for known token formats

organization-specific patterns for internal keys and connection strings

checksums and validators for structured identifiers

named-entity recognition for personal information

dictionaries for internal project names or classified terms

entropy and context signals for unknown high-randomness secrets

allowlists for known false positives

human review for high-impact ambiguous cases

Detection systems produce false positives and false negatives. The answer is not to disable them. Tune them with labeled examples from the organization’s real data and measure both precision and recall.

Add a Sanitization Boundary in Code

The following simplified Python example shows an application-side boundary. It blocks credentials, replaces selected confidential values with typed placeholders, and returns findings separately from sanitized content. It intentionally does not log the original text.

In production, replace the example patterns with a supported DLP or PII detector, organization-specific recognizers, and a policy service. Keep the interface stable so the application does not need to know which detection engine is underneath it.

from __future__ import annotations

import re
from dataclasses import dataclass
from enum import StrEnum
from typing import Iterable

class DataClass(StrEnum):
CONFIDENTIAL = “confidential”
CREDENTIAL = “credential”

@dataclass(frozen=True)
class Finding:
name: str
data_class: DataClass
start: int
end: int

class SensitiveDataBlocked(RuntimeError):
pass

DETECTORS: tuple[tuple[str, DataClass, re.Pattern[str]], …] = (
(
“private_key_material”,
DataClass.CREDENTIAL,
re.compile(r”—–BEGIN [A-Z ]+PRIVATE KEY—–“),
),
(
“generic_bearer_value”,
DataClass.CREDENTIAL,
re.compile(r”(?i)bbearers+[a-z0-9._~+/=-]{20,}”),
),
(
“email_address”,
DataClass.CONFIDENTIAL,
re.compile(r”b[A-Z0-9._%+-]+@[A-Z0-9.-]+.[A-Z]{2,}b”, re.I),
),
)

def detect(text: str) -> list[Finding]:
findings: list[Finding] = []
for name, data_class, pattern in DETECTORS:
for match in pattern.finditer(text):
findings.append(
Finding(
name=name,
data_class=data_class,
start=match.start(),
end=match.end(),
)
)
return sorted(findings, key=lambda item: item.start)

def sanitize(text: str, findings: Iterable[Finding]) -> str:
findings_list = list(findings)

if any(item.data_class == DataClass.CREDENTIAL for item in findings_list):
raise SensitiveDataBlocked(“Credential-like content was blocked”)

output = text
for item in sorted(findings_list, key=lambda value: value.start, reverse=True):
replacement = f”[{item.name.upper()}]”
output = output[: item.start] + replacement + output[item.end :]

return output

def prepare_model_input(raw_text: str) -> tuple[str, list[str]]:
findings = detect(raw_text)
sanitized = sanitize(raw_text, findings)
finding_names = sorted({item.name for item in findings})
return sanitized, finding_names

The application should record that a credential detector blocked the request, along with a policy ID, request ID, environment, and detector version. It should not record the matched value or the surrounding raw text.

What can go wrong:

overlapping findings can create replacement errors unless the library handles spans correctly

broad patterns can block legitimate content

encoded or split secrets can bypass simple matching

language-specific PII can be missed

model-generated output can reintroduce sensitive data after input redaction

For those reasons, use the code pattern as an architectural boundary, not as a complete enterprise DLP implementation.

Protect Embeddings and Vector Stores as Sensitive Data Systems

Embeddings are not a privacy eraser. They are derived from source data and can preserve information that should remain protected. The vector store also usually contains chunk text, document metadata, source identifiers, access labels, and retrieval history.

Treat the retrieval layer as a data platform with the same rigor as a database.

Preserve Source Authorization

Every indexed chunk should retain enough metadata to enforce the source system’s permissions at query time. At minimum, include:

source system and object identifier

tenant or business boundary

data classification

owner

allowed groups, roles, or policy reference

ingestion timestamp and pipeline version

source version or content hash

retention and deletion metadata

Do not rely on the model to ignore a retrieved chunk that the user was not authorized to see. Unauthorized content should never enter the model context.

Partition High-Risk Data

Shared indexes increase the blast radius of a permission error. Use separate collections, namespaces, projects, accounts, or clusters when tenant, regulatory, geographic, or administrative boundaries require stronger isolation.

Logical filters are useful, but they are not always an adequate substitute for physical or administrative separation. The right boundary depends on the impact of a missed filter and the ability to prove isolation.

Design Deletion Before Ingestion

A source-record deletion must propagate through:

extracted text

chunk stores

embeddings

vector indexes

replicas

caches

evaluation sets

snapshots and backups

Store a stable lineage key so the platform can find every derived object for a source record. If the only way to delete one record is to rebuild the entire index, document that behavior and make the rebuild process repeatable.

Make Observability Useful Without Capturing Content

AI observability does not require full prompt capture. Operators can diagnose most availability, latency, model-routing, token, retry, and tool-performance problems with metadata.

A production-safe telemetry contract should prefer:

service, environment, agent, workflow, and prompt version

requested and response model identifiers

input and output token counts

latency and time-to-first-response

finish reason

tool name and status

retry count

policy decision and detector category

low-cardinality error type

trace and request correlation identifiers

It should exclude by default:

system instructions

user prompts

model responses

retrieved documents

tool arguments and results

authorization headers

cookies and session tokens

database queries with parameter values

full exception messages and stack-local values

user identifiers as metric dimensions

Filter at the Application and Collector

Filtering only at the telemetry backend is too late. Raw content may already have crossed networks, entered queues, or been retained by an agent or collector.

Use at least two enforcement points:

Application instrumentation: do not create sensitive attributes.

Collector or gateway: delete, hash, or transform prohibited fields before export.

Backend controls are still useful as a final guardrail, but they should not be the first place sensitive content is removed.

Treat Debug Capture as a Controlled Exception

There are cases where short-lived content capture is necessary to diagnose a production failure. That should be an exception workflow, not a permanent feature flag.

A controlled debug session should define:

incident or change ticket

named owner and approver

exact service and environment

approved data class

start and automatic expiration time

restricted destination

reduced sampling rate

deletion owner

validation that capture stopped and data was purged

Prefer synthetic or replayed sanitized requests before enabling production content capture.

Scan Source Control, Prompts, and Build Artifacts for Secrets

Secrets do not enter AI systems only through user prompts. They also appear in:

prompt templates

notebooks

evaluation examples

test fixtures

configuration files

infrastructure-as-code variables

CI/CD logs

generated documentation

pull request descriptions and comments

issue trackers

exported traces and support bundles

Enable secret scanning across the full repository history, not only the latest branch. Add push protection or an equivalent pre-receive control so known credentials are blocked before they become durable history.

Use custom patterns for organization-specific tokens and internal connection strings. Validate findings when the scanning platform supports it, but do not delay revocation of a confirmed credential while waiting for perfect classification.

A leaked secret should be rotated or revoked. Deleting the line from the repository does not make a credential safe again.

Set Retention by Surface, Not by Application

One “AI retention period” is rarely sufficient because different data surfaces serve different purposes.

Use a retention matrix such as the following as a design starting point. Final values must align with records, privacy, security, operational, and legal requirements.

Data surface
Recommended default posture
Required deletion trigger

Metrics without content
Retain long enough for capacity and trend analysis
End of approved operational period

Traces without prompt or tool content
Short operational window
Age-based deletion or incident closure

Debug traces with approved content
Hours or a few days, not routine retention
Automatic expiration plus verified purge

Application logs
Structured metadata only, shortest useful period
Age, incident closure, or policy change

Conversation memory
User- and use-case-specific
Session expiration, user request, account closure

Vector chunks and embeddings
Match source authorization and lifecycle
Source deletion, permission change, dataset retirement

Evaluation sets
Sanitized and versioned
Test retirement, consent or source deletion obligation

Security audit evidence
Governed security retention
Policy-defined expiration or legal hold release

Backups and snapshots
Encrypted and time-bounded
Backup lifecycle expiration and deletion verification

Retention should be enforceable through lifecycle policies, not calendar reminders. Every exception should expire automatically.

Verify Deletion Across Replicas

Deletion proof should answer:

Was the primary record deleted?

Was the vector entry removed or the index rebuilt?

Were caches invalidated?

Were search replicas updated?

Were evaluation copies removed?

Will backups expire within the approved period?

Can the data still be retrieved by semantic search?

Can an observability query still find the value?

A successful API deletion response is not sufficient evidence when the workflow created multiple derived copies.

Encrypt Every Persistent Store and Separate Key Ownership

Encryption does not replace redaction or access control, but it reduces the impact of storage theft, snapshot exposure, and misplaced exports.

Protect:

prompt and conversation databases

vector databases and chunk stores

object storage used for ingestion

logs, traces, and dead-letter queues

evaluation and experiment platforms

caches that persist to disk

backups, snapshots, and exports

Use encryption in transit and at rest. Prefer customer-managed keys or equivalent controls when risk, regulation, or operational ownership requires them. Separate key administrators from data administrators where practical, and restrict export, snapshot, and restore permissions.

Service identities should receive only the permissions required for their role. The embedding worker should not automatically have broad read access to every document collection. The telemetry collector should not be able to decrypt restricted source data. The model runtime should not be able to enumerate secrets.

Run Validation Before Production and After Every Material Change

Sensitive-data controls need regression tests just like application logic. Build a synthetic corpus containing fake credentials, personal information, proprietary terms, encoded values, and known false positives.

Admission-Control Tests

Confirm that:

credential-like values are blocked

restricted records require an approved exception

confidential values are transformed as expected

allowed public and internal content still works

blocked content does not appear in error logs or retry queues

Retrieval Tests

Confirm that:

a user cannot retrieve another tenant’s chunks

permission changes take effect without waiting for a full reindex when the architecture promises that behavior

deleted records cannot be recovered through semantic search

poisoned or untrusted documents are quarantined before indexing

source metadata and lineage remain intact

Telemetry Tests

Confirm that:

prompt and response content are absent by default

tool arguments and results are absent by default

authorization headers and cookies are removed

metric dimensions do not include user data or unbounded text

collector-side filters still work if an application accidentally emits a prohibited field

debug capture expires automatically

Secret-Scanning Tests

Commit synthetic test patterns in an isolated test repository and verify that scanning and push protection detect them. Do not use active credentials for control testing.

Release Gate

Fail the release when any synthetic restricted value reaches a model request, embedding record, trace, log, evaluation set, or export without the expected policy action.

Contain a Sensitive Data Leak as a Multi-Surface Incident

When sensitive content is found, do not begin by deleting the first log entry. First contain active exposure and preserve enough non-sensitive evidence to determine scope.

The decision flow below separates credential response from non-credential data response while preserving a common containment and validation path.

Containment Actions

Use the following sequence:

Classify the exposed value. Determine whether it is a credential, personal data, proprietary information, regulated content, or a combination.

Revoke or rotate active credentials immediately. Assume copying occurred even when access logs do not show misuse.

Disable content-bearing telemetry. Turn off prompt, response, tool-argument, and tool-result capture at the application and collector.

Pause affected ingestion. Stop new documents from entering the contaminated pipeline.

Restrict access. Reduce access to affected dashboards, vector collections, buckets, queues, and exports.

Quarantine affected indexes or data sets. Prevent further retrieval while scope is determined.

Preserve safe evidence. Retain identifiers, timestamps, hashes, policy versions, and access events without copying the sensitive value into the incident record.

Scope the Replication Path

Search every system that could have received the value:

original request and conversation store

model gateway and provider request records

application logs

traces and events

message queues and dead-letter queues

caches

vector chunks, embeddings, and metadata

evaluation and feedback platforms

code repositories and CI/CD artifacts

support tickets and chat channels

exports, snapshots, and backups

Search with a cryptographic hash, stable fingerprint, source record ID, or tightly restricted exact-match procedure when possible. Avoid copying the raw value into general-purpose search tools.

Eradication and Recovery

Depending on the surface, remediation may require:

deleting a prompt or conversation record

rebuilding a vector index from a clean source set

purging caches

deleting or expiring trace partitions

replacing contaminated evaluation cases

revoking shared links and exports

rotating keys and tokens

replaying sanitized transactions

updating detectors and policy rules

Restore service only after a validation test confirms that the value cannot be retrieved, reproduced, or observed through the affected paths.

Notification and Post-Incident Work

Security, privacy, legal, records management, data owners, application owners, and third-party providers may all need to participate. The incident process should determine notification obligations based on the data, affected people or systems, processing locations, contracts, and applicable requirements.

The post-incident review should produce control changes, not only awareness reminders. Ask which automated boundary failed, why the data was replicated, why retention allowed it to persist, and which test should prevent recurrence.

Common Failure Modes

Relying on a System Prompt to Protect Secrets

A prompt instruction is not deterministic access control. Secrets and authorization logic should remain outside model-visible context.

Redacting the Prompt but Logging the Original

The blocked or transformed request often appears in debug messages, exception objects, queue payloads, and trace events. Inspect the failure path, not only the successful provider call.

Treating Embeddings as Anonymous

Derived data can still be sensitive, and vector stores commonly retain source text and metadata. Apply classification, access control, encryption, lineage, and deletion.

Capturing Everything for Observability

Full content may make one debugging session easier while creating a long-lived privacy and security problem. Start with metadata, then enable narrowly scoped content capture only through an expiring exception.

Using DLP as the Only Control

Detectors miss context, encoded content, novel secret formats, and business-sensitive information that has no obvious pattern. Combine DLP with source authorization, minimization, isolation, output filtering, retention, and incident response.

Deleting Only the Primary Record

AI workflows create derived and replicated data. A complete deletion must include chunks, embeddings, caches, logs, evaluations, snapshots, and exports according to their approved lifecycle.

Keeping Permanent Exceptions

A temporary diagnostic flag becomes a permanent data-collection path unless it has automatic expiration. Exceptions need an owner, reason, scope, and end time.

Operational Ownership

Sensitive-data protection fails when every team owns one component but nobody owns the complete path.

Control area
Primary owner
Required partners

Classification and approved use
Data owner
Privacy, security, legal, records management

Prompt and input inspection
Application team
AI platform, security engineering

DLP and redaction service
Security or platform team
Application teams, data owners

Vector authorization and lifecycle
AI platform or data platform
Source-system owners, identity team

Telemetry contract and filtering
Observability platform
Application teams, security, privacy

Secrets management and rotation
Security platform
Application and infrastructure teams

Retention and deletion
Records and data governance
Platform owners, privacy, legal

Incident containment
Security incident response
All affected system and data owners

One role should be accountable for proving the end-to-end control, even when implementation spans several platforms.

Production Readiness Checklist

Before enabling production traffic, confirm:

data classes and default AI actions are documented

credentials and restricted data are denied by default

prompt assembly has an inspection and transformation boundary

secrets are externalized from prompts and model-visible tool data

retrieval preserves source permissions and tenant boundaries

embeddings, chunks, metadata, replicas, and backups have deletion paths

prompt, response, and tool-content telemetry are disabled by default

collector-side filters remove prohibited fields

debug capture requires approval and expires automatically

source control, prompt assets, notebooks, and CI/CD outputs are secret-scanned

every persistent store is encrypted and access-reviewed

retention is defined separately for each data surface

synthetic leak tests run in CI/CD and release validation

credential rotation and multi-surface containment procedures are tested

ownership and escalation paths are current

Conclusion

Keeping secrets and sensitive data out of AI systems requires more than asking users to be careful. The control must exist in the architecture.

Classify data before it enters the workflow. Block credentials and restricted content by default. Redact, tokenize, or minimize confidential data only when the use case is approved. Preserve source authorization through retrieval. Treat embeddings and vector stores as sensitive data systems. Capture observability metadata without collecting raw content. Scan code and prompt assets before they become durable history. Encrypt persistent stores, define retention by surface, and design deletion before ingestion.

When exposure occurs, assume replication. Rotate active credentials, stop capture, quarantine affected data sets, trace every derived copy, purge or rebuild contaminated stores, and prove that the value cannot be recovered before restoring normal operation.

The practical security boundary is not the model. It is the set of deterministic controls around the complete AI data path.

External References

National Institute of Standards and Technology: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileCanonical URL: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence

National Institute of Standards and Technology: Guide to Computer Security Log ManagementCanonical URL: https://csrc.nist.gov/pubs/sp/800/92/final

OWASP Generative AI Security Project: LLM02:2025 Sensitive Information DisclosureCanonical URL: https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/

OWASP Generative AI Security Project: LLM07:2025 System Prompt LeakageCanonical URL: https://genai.owasp.org/llmrisk/llm072025-system-prompt-leakage/

OWASP Generative AI Security Project: LLM08:2025 Vector and Embedding WeaknessesCanonical URL: https://genai.owasp.org/llmrisk/llm082025-vector-and-embedding-weaknesses/

OpenTelemetry: Inside the LLM Call: GenAI Observability with OpenTelemetryCanonical URL: https://opentelemetry.io/blog/2026/genai-observability/

GitHub Docs: Secret scanningCanonical URL: https://docs.github.com/en/code-security/concepts/secret-security/secret-scanning

Google Cloud Documentation: De-identifying sensitive dataCanonical URL: https://docs.cloud.google.com/sensitive-data-protection/docs/deidentify-sensitive-data

Presidio: Presidio AnalyzerCanonical URL: https://presidio.dataprivacystack.org/analyzer/

Presidio: Presidio AnonymizerCanonical URL: https://presidio.dataprivacystack.org/anonymizer/

Operationalizing the Agent Blast Radius Model: Policy Gates, Tool Contracts, and Rollback Controls
A risk model is only useful if it changes what happens before the agent acts. The Agent Blast Radius Model gives teams…

Next PostWho Owns the Failure? Building a Support RACI for a Multivendor Private AI PlatformTL;DR A multivendor private AI platform is not operationally complete when the hardware is installed, the GPUs are visible, and the first model endpoint responds. It is complete when the…
The post How to Keep Secrets and Sensitive Data Out of AI Prompts, Embeddings, Traces, and Logs appeared first on Digital Thought Disruption.