Skip to content
Magic

Ingest.Extract.Search.

DocSlurp transforms any document into structured, searchable intelligence. Smart LLM routing, table extraction, image captioning, citations — all through one API.

quickstart.ts
import { createDocSlurpClient } from '@docslurp/sdk';
import { readFile } from 'node:fs/promises';

const client = createDocSlurpClient({
  apiUrl: process.env.DOCSLURP_URL!,
  apiKey: process.env.DOCSLURP_API_KEY,
});
const { workspace, run } = await client.upload({
  files: [{ filename: 'report.pdf',
    data: new Blob([await readFile('report.pdf')]) }],
});
await client.waitForRun(run.id);
const answer = await client.chat({
  workspaceId: workspace.id,
  q: 'What are the key findings?',
});

PDF to structured data
in one API call.

From raw PDFs to structured, searchable intelligence — DocSlurp handles the entire pipeline.

Dynamic LLM Routing

Auto-selects the optimal model per document region — text, tables, images, mixed content.

Advanced Table Extraction

Variable columns, merged cells, mixed content types — all handled with structural fidelity.

Image Intelligence

Turn document images into searchable context alongside extracted text.

Citations & Page Images

Every search result links back to exact page regions with rendered page image previews.

Smart Region Detection

Identifies headers, footers, sidebars, figures, and body text with spatial precision.

Enrichment Pipeline

Entity extraction, sentiment analysis, topic classification — layered onto every document.

Interactive Word Clouds

Dynamic, filterable word cloud visualizations for rapid corpus understanding.

Subject & Intent Graphs

Explore relationships between topics, entities, and document intent in interactive graphs.

Optimized Processing

Parallel extraction pipelines with intelligent caching for sub-second re-queries.

Focused Retrieval

Search within a workspace and narrow results by document, filename, or tags.

Agent Tools

Authenticated REST tools for workspace search and grounded agent integrations.

SDK, API & CLI

TypeScript SDK, REST API, and CLI — pick your interface, ship in minutes.

The pipeline

From upload to searchable documents, with sensible defaults.

Step 1

Ingest

Upload PDFs, DOCX, images, HTML — any document format.

Format detectionParallel chunkingPage rendering
Step 2

Process

LLM-powered extraction with smart model routing.

Region detectionTable parsingImage captioning
Step 3

Enrich

Entity extraction, classification, and relationship mapping.

NER & sentimentTopic graphsWord clouds
Step 4

Search

Semantic search with citations, page images, and filters.

Vector + keywordCited resultsFaceted filters

Three interfaces. One platform.

SDK for developers, REST for integrations, tools for AI agents — choose your path.

SDK
TypeScript SDK

Upload, wait for processing, search, and ask questions with a small typed client.

import { createDocSlurpClient } from '@docslurp/sdk';

const client = createDocSlurpClient({
  apiUrl: process.env.DOCSLURP_URL!,
  apiKey: process.env.DOCSLURP_API_KEY,
});

const { results } = await client.search({
  workspaceId: 'workspace-id',
  q: 'revenue Q4',
});
for (const hit of results) {
  console.log(hit.text, hit.citation);
}
API
REST API

Upload batches, track processing, and retrieve cited answers through REST.

curl -X POST "$DOCSLURP_URL/v1/ingest/bulk" \
  -H "Authorization: Bearer $DOCSLURP_API_KEY" \
  -F "[email protected]"

# Returns workspace, run, and documents.
# Check /v1/runs/<run-id>/progress
# before searching your workspace.
AGENTS
Agent Tools

Authenticated REST tools let your agents search the same workspace documents.

curl -X POST \
  "$DOCSLURP_URL/v1/mcp/tools/workspace.search" \
  -H "Authorization: Bearer $DOCSLURP_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"workspaceId":"workspace-id",
       "q":"revenue Q4"}'

# REST tools for agent integrations.

Start building in minutes.

Drop files, auto-create workspaces, and surface grounded search — no configuration required.