Powered by Small Language Models.

Unlike LLM-based security tools that send your data to external APIs, Polygraf runs specialized SLMs entirely on your infrastructure — delivering enterprise-grade accuracy with zero data exposure.

Why SLMs Over LLMs.

Privacy by Design

Every model runs locally on your hardware. No data is ever transmitted to external services. Fully air-gap compatible.

0 bytes sent externally

Sub-100ms Latency

Specialized models are orders of magnitude smaller and faster than general-purpose LLMs. Users experience zero friction.

<100ms average

Domain Precision

Each SLM is purpose-trained for a specific security task — PII detection, AI provenance, credential scanning — with higher accuracy than generalist models.

93–99% accuracy

Comparison

SLM vs LLM vs Rule-Based

Based on internal benchmarking across standard enterprise workloads. LLM latency figures from published API documentation. Rule-based false positive rates from industry research (Gartner, 2023).

The Reality

AI adoption is outpacing security by years.

89%

of employees use AI tools not approved by IT

1 in 3

AI queries contains sensitive data — PII, credentials, or IP

0

traditional DLP tools were built to handle LLM interactions

98%

Detection accuracy

across all entity types

<100ms

Max latency

imperceptible to users

35+

Entity types

PII, PHI, credentials, IP

100%

On-premises

PII, PHI, credentials, IP

49+ Entity Types.

Out-of-the-box detection across every sensitive data category — plus custom entity training for your organization's unique identifiers.

Personal Identity

Contact & Location
Financial
Credentials & Access
Organization & Work

Contextual & Misc

Performance Benchmarks.

Macro F1 and Weighted F1 scores across industry-standard NER datasets. Polygraf consistently outperforms cloud-based competitors — while keeping all data on-premises.

Technical Questions.

What exactly is a Small Language Model (SLM)?

An SLM is a neural network model with tens to hundreds of millions of parameters — purpose-trained for a narrow task. Unlike general-purpose LLMs (which have billions of parameters and broad capabilities), Polygraf’s SLMs are each trained on a specific security task like PII detection or credential scanning. This makes them much faster, cheaper to run, and more accurate for that specific task.

For the specific task of sensitive data detection, our SLMs outperform general-purpose LLMs on every standard benchmark (ONTONOTES, GMB, RE3D, and our proprietary dataset). GPT-4 and Claude are excellent at broad reasoning, but they were not designed for this task and carry significantly higher false-positive and false-negative rates.

Yes. Polygraf supports custom entity training. You provide labeled examples of your organization-specific identifiers (e.g. internal project codes, employee IDs, proprietary product names), and we fine-tune a dedicated SLM for your environment.

Most organizations run Polygraf on standard enterprise server hardware. A recommended baseline is a 16-core CPU server with 64GB RAM. GPU acceleration is supported but not required. Models are quantized for efficient inference.

Macro F1 averages the F1 score across all entity classes equally, regardless of class size — it’s a more stringent measure of broad coverage. Weighted F1 weights each class by its frequency in the dataset. Both are standard NLP evaluation metrics. We report both for full transparency.

The benchmarks shown are from our internal evaluation lab using publicly available datasets (ONTONOTES, GMB, RE3D). We are in the process of engaging a third-party auditor for independent verification. Full methodology and evaluation scripts are available under NDA.

Get the Full Technical Brief.

Deep dive into SLM architecture, benchmarks, and deployment specs

Products

thank you

Your download will start now.

Thank you!

Please provide information below and
we will send you a link to download the white paper.