Data Discovery
Find sensitive data everywhere it lives, including the places nobody remembers creating.
The problem
Most organizations can name their crown-jewel databases. Almost none can name the export in a personal drive, the customer list pasted into a Slack thread in 2023, or the staging bucket a contractor filled and forgot. Discovery closes the gap between the data you govern and the data you actually have.
How it works
What it covers
- Transformer-based semantic classification
- Named-entity recognition tuned per data type
- Validated pattern detection for structured identifiers
- File-type aware parsing (documents, spreadsheets, archives, images via OCR)
- Continuous sensitive-data inventory with owners and locations
- Exposure flags for public links, external shares, and stale access
- Data maps exportable to GRC and privacy tooling
- Feeds classification, lineage, DSPM, and DLP with the same findings
- Discovers sensitive data inside prompt logs and AI tool exports
- Inventories vector stores and embedding sources
- Flags datasets staged for fine-tuning that contain regulated data
Common questions
Does discovery require agents everywhere?
No. SaaS, cloud, and data-platform coverage is API-based. The endpoint agent is only needed for device-level visibility and is optional per fleet.
How fresh is the inventory?
Connectors stream change events where the source supports them and poll incrementally where it does not. Most environments stay within minutes of reality.
Will scanning move my data out of region?
With the private data plane option, content is scanned inside your boundary and only metadata reaches the control plane.
Solutions this powers
See Data Discovery on your data.
Connect one environment in a guided session and review real findings with a security engineer.