DataZen

The Integration Enginefor the AI Era

One engine to connect business systems, process data as it moves, and give AI governed access to approved tools and workflows.

Access governed DataZen pipelines securely from the AI assistants you already use.

CopilotClaudeChatGPTGeminiCursor

Trusted Clients and Partners

A SINGLE INTEGRATION LAYER

One engine, from source to delivery.

DataZen integration engine Acquire data, transform, enhance and validate it, then route and deliver. Parallel connections send and receive data through Cloud Functions and Agentic AI.

Connect, Transform, Enrich, Deliver, and Mobilize Data in one continuous pipeline.

See DataZen in action

Connect to Anything

Read from APIs, databases, files, and messaging systems with one SQL-first engine. Start from a source, not from a blank integration project.

Query and Combine Anything

Join live API responses the same way you join tables. Combine records across systems without extra staging layers or brittle custom glue code.

Put AI Inside the Pipeline

Use AI inline to classify, normalize, summarize, and enrich data as it moves. Keep prompts inside the pipeline instead of outside the workflow.

Turn Pipelines into Governed AI Tools

Turn approved pipelines into governed tools for AI assistants. Expose only the right capability, with controlled access, auditability, and secure execution.

THE DATAZEN ARCHITECTURE

Built differently from traditional integration platforms.

Five Architectural Differences that make DataZen simple and flexible.

DataZen

Universal
Connectivity

DataZen

Inline AI

DataZen

Synthetic
CDC

DataZen

Zero-Landing
Architecture

DataZen

Pipelines as
Tools

WHY THE ARCHITECTURE MATTERS

Common Approach vs. DataZen

SECURITY BY ARCHITECTURE

Turn your integration layer into a governed security boundary

DataZen sits between the tools that initiate work and the systems that hold your data. Credentials stay centralized, connections are explicitly approved, and permissions stay narrowly scoped.

That becomes especially important as AI enters the workflow. An AI tool gets only the access assigned to that tool, not everything available to the person using it. You decide which systems it can reach, what data it can access, and what actions it can perform.

KEY CONTROLS

Centralized credentials

Keep connection secrets in DataZen instead of distributing them across users, scripts, applications, or AI tools.

Least-privilege access

Limit connections to the systems, tools, and operations they actually need.

Transfer & access logging

Keep a record of pipeline activity, transfers, and access across integrations.

Approved connections

Control which sources and destinations a workflow is allowed to reach.

Flexible deployment

Run on-premises, in a private cloud, Cloud VPC, hybrid environment, or isolated infrastructure.

Users
Workflows
AI assistants
Client apps
External services
DataZen ToolGateGoverned integration control layer
Centralized credentials
Scoped access
Approved connections
Transfer & access logs
Flexible deployment

Your environment

Applications
Databases
Files
Internal APIs
Pricing

Start small. Scale as pipelines grow

Choose the DataZen tier that matches your current integration needs, then expand as sources, pipelines, and execution demands increase.

Bronze
Entry plan for core data acquisition, database targets, and starter pipeline usage.
$99 /mo
  • HTTP, DB*, DRIVE, and BIG DATA sources
  • Database targets
  • ETL / ELT included
  • 50 active pipelines
  • 1 day log history retention
Connections10
Pipelines50
Runtime5 min
Executions250/day
Silver
Expanded sources and targets with CDC, replay, API access, and AI-enabled pipeline tools.
$499 /mo
  • HTTP, DB, DRIVE, and BIG DATA targets
  • Change Data Capture
  • Replay and X.509 encryption
  • Limited DataZen API access
  • 7 days log history retention
Connections25
Pipelines250
Runtime15 min
Executions1000/day
Gold
Highest-scale plan for broad connectivity, enterprise limits, and advanced operations.
Contact Us
  • Any source, any target
  • Messaging and ODBC support
  • Unlimited executions and compute
  • 1 minute scheduling frequency
  • 30 days log history retention
Connections100+
Pipelines1000+
RuntimeUnlimited
ExecutionsUnlimited
Contact Us

Compare Plan Features

Sources & Targets

Bronze
Silver
Gold
Sources
HTTP, DB*, DRIVE, BIG DATA
HTTP, DB*, DRIVE, BIG DATA
HTTP, DB, DRIVE, BIG DATA, MESSAGING, ODBC
Targets
DB
HTTP, DB, DRIVE, BIG DATA
HTTP, DB, DRIVE, MESSAGING, BIG DATA, ODBC

Pipelines

Bronze
Silver
Gold
Active Pipelines
50
250
1000+
ETL / ELT
Included
Included
Included
Env Variables
Included **
Included
Included
CDC
—
Included
Included
Replay
—
Included
Included
X.509 Encryption
—
Included
Included
DataZen API
—
Limited ***
Available
Retention
1 day
7 days
30 days

Limits

Bronze
Silver
Gold
Connections
10
25
100+
Max Storage
5GB
20GB
100GB
Daily Executions
250 incl. preview operations
1,000 incl. preview operations
Unlimited
Run Duration
5 min / 1 hour
15 min / 2 hours
Unlimited
Schedule Frequency
1 hour
15 min
1 min
Compute
10 min / 1 hour
30 min / 1 hour
Unlimited
Max Data Transfer / Month
2 GB
25 GB
100 GB+
Peak Memory
50MB / 250MB
250MB / 500MB
2GB / 2GB

Other

Bronze
Silver
Gold
Portal User Mgmt
—
Included
Included
Security Admin
—
Limited ***
Included
IP Firewall
Included
Included
Included
AI Chat
—
Included
Included
MCP
—
Included
Included

* DB includes SQL Server, MySQL, Postgres, Snowflake, and Fabric. ODBC is available in Gold only.

** Environment variable encryption is available in Silver edition or higher.

*** Essential access includes creating an API key, starting and stopping pipelines, and obtaining pipeline summaries and status information for agentic skills, integration, and orchestration.

Pricing is per agent per month, paid upfront, and is non-refundable; additional consumption charges may apply, including storage and compute overages.

Supported database targets include SQL Server, MySQL, Postgres, Snowflake, and Fabric. Self-hosted agent priced separately — contact us for details.

DATAZEN FAQ

Frequently asked questions

The Basics

What exactly is DataZen?

DataZen is an integration engine for building data pipelines across APIs, databases, files, cloud services, applications, messaging systems, and AI. It offers key capabilities designed for rapid integration and data movement, including:

  • In-memory data processing and inline ETL
  • Data engineering features, such as change data capture, high-watermark processing, ETL commands, and more
  • Limited storage requirements to minimize the need for Medallion Bronze and Silver layers
  • Exposing and governing business data pipelines as MCP Tools for popular AI Agents
  • Engaging AI Agents and HTTP/S endpoints inline as part of the ETL pipeline

It reads data from where it already lives using familiar SQL-like syntax, transforms and enriches it as it moves, identifies meaningful changes when needed, and delivers the result wherever it needs to go.

The same platform can support ETL/ELT, application integration, synchronization, replication, automation, CDC, and AI-powered workflows.

What makes DataZen so different from traditional integration platforms?

In-Memory

A foundational difference is that DataZen can run a complete pipeline in memory without requiring data to be landed or staged between steps in most cases. Traditional integration platforms typically break processing into separate stages: extract, land, stage, transform, hand off, and only then deliver the data to its final destination. DataZen can keep all this work moving through a single pipeline, reducing the need for intermediate layers.

Inline Data Enrichment

What further sets DataZen apart is its ability to enhance or modify data by calling other endpoints, HTTP/S APIs, AI Agents, and databases as part of the pipeline itself. Because the data remains entirely in memory, these calls can all happen within the same continuous pipeline.

Switching Language In-Flight

Traditional integration stacks put much of their complexity into the integration itself, requiring different technologies, logic, and processing to be coordinated across multiple stages, tools, and environments.

DataZen can switch between execution contexts inline, from in-memory processing to HTTP/S APIs to databases and back again. This makes it possible to build heterogeneous pipelines that stitch together multiple programming paradigms in one continuous flow, including custom APIs built on-premises and cloud functions such as AWS Lambda in Python or Azure Functions in C#.

The bottom line is that you may still use warehouses, lakes, staging systems, and other integration tools where they make sense. But DataZen simply does not require those layers to sit in the middle of every integration.

Can DataZen work with the infrastructure we already use?

Yes.

DataZen is designed to work alongside existing databases, APIs, applications, warehouses, lakes, cloud services, and integration systems.

You can introduce it at specific points in your existing architecture—at the source, around Bronze or Silver layers, or for individual operational integrations—without redesigning the entire data estate.

You can also rebuild your existing pipelines from scratch.

Do we have to replace our current integration platform?

No.

DataZen can augment an existing integration stack rather than replace it all at once. You can keep the systems and platforms that already work and use DataZen where it simplifies an integration, removes unnecessary layers, adds CDC or enrichment, connects a difficult source, or adds capabilities you don't currently have.

Adoption can happen pipeline by pipeline.

Moving and Processing Data

Do I have to land data somewhere before I can transform it?

No.

DataZen can transform, filter, enrich, mask, validate, and otherwise work with data inline while it moves through the pipeline.

You can still persist raw or intermediate data when there is a reason to, such as for history, compliance, analytics, auditing, or replay. The difference is that, in DataZen, a permanent landing layer is not a required step in every integration.

Do I still need Bronze and Silver layers?

Sometimes, but not automatically.

If your architecture benefits from a raw historical Bronze layer or persistent Silver data, DataZen can readily work alongside those layers and seamlessly incorporate them into the pipeline. This architectural decision can depend on compliance requirements, the performance of other systems, or long-running transaction tracking.

But DataZen can also drastically reduce storage requirements by transforming and enriching data inline, using transient database processing when needed, and then sending your business-ready data directly to its destination. That means Bronze and Silver become architectural choices rather than mandatory steps imposed by the integration platform.

Another way DataZen significantly reduces storage costs is through its change-log model. DataZen can capture the last known state of your data in compressed, schema-aware change logs that hold either an entire dataset or only the changes since the last read. Because DataZen creates these logs only when needed, it avoids storing unnecessary copies of data. At the same time, the logs provide many of the functions traditionally handled by Bronze storage layers, including persistence, recovery, and replay. Their compact, portable format makes it easy to reconstruct data states, rerun processing, and move changes between environments.

Can DataZen use the compute power of databases we already have?

Yes.

Not every transformation has to run inside DataZen's in-memory engine.

For work such as complex joins, aggregations, multi-step SQL, or operations against large reference datasets, DataZen can move the current pipeline dataset into transient database tables, run database-native SQL, return the result to the pipeline, and clean up those temporary tables when the block completes.

This makes the choice of processing environment a design decision at each step. As you build the pipeline, you can switch to a different processing environment as needed, at any step along the way. For example, one step can use DataZen's in-memory engine, and the next can use database-native processing, all within the same pipeline. This also lets you route work to lower-cost processing resources where it makes sense, rather than relying on the same processing environment for every step.

Change Data Capture

How does DataZen avoid processing the same unchanged data over and over?

DataZen can use Synthetic CDC to identify which records have actually changed between executions, even when the source doesn't provide a traditional change stream. Synthetic CDC is a high-performance in-memory operation that works on any data source, including files, HTTP/S endpoints, and databases.

That lets downstream processing focus on inserts, updates, and deletions instead of repeatedly processing records whose business state hasn't changed.

When a source provides a reliable watermark, DataZen can also limit how much data needs to be retrieved in the first place. These techniques can be combined to both reduce data retrieval and eliminate unnecessary processing.

Can the same pipeline send data to more than one destination?

Yes.

A DataZen Reader can capture the data once and allow multiple Writers to consume the same Change Logs independently. DataZen refers to this pattern as multicasting.

That means a single captured change set can feed several applications, databases, warehouses, lakes, or other targets without independently recapturing data from the source for each destination.

Connecting Systems

What kinds of systems can DataZen connect to?

An unlimited number.

DataZen supports connections to HTTP/S APIs, relational databases via ODBC drivers, cloud drives and file systems, big-data platforms, messaging systems, and other enterprise sources and targets.

Available connection types may vary between cloud-hosted and self-hosted deployments.

Do we need a prebuilt connector for every system we want to use?

No.

DataZen can connect directly to HTTP/S endpoints, databases, files, storage systems, messaging technologies, and other supported interfaces.

For APIs, DataZen can work with OpenAPI, Swagger, and Postman specifications, removing much of the dependency on vendor-built connectors. Rather than waiting for a dedicated connector to be built or updated for an application, DataZen can communicate directly with systems where accessible APIs and endpoints are available.

Unlike other integration platforms, DataZen’s HTTP/S engine is not limited by a catalog of prebuilt connectors. If DataZen can connect, it can extract and/or send data.

DataZen supports a broad range of enterprise authentication methods, including Basic Auth, Bearer Tokens, OAuth 2.0, session-based authentication including JWT, and Negotiated authentication for on-premise deployments.

What if we need to connect to a custom or proprietary API?

That's exactly the kind of situation where DataZen's HTTP connectivity is useful.

You can work with internal APIs, proprietary applications, industry-specific services, cloud functions, and custom HTTP endpoints without waiting for a product-specific connector to be created.

Connections can be configured once and reused across multiple pipelines.

How does DataZen handle a schema change?

DataZen includes schema-management operations for inspecting, reshaping, enforcing, and adapting schemas inside a pipeline.

For supported database targets, DataZen can also automatically create merge logic and apply schema drift, including adjustments to columns and data types. Exact schema-drift behavior varies by database engine.

That gives you control over how schema changes should flow through an integration.

Can DataZen call another API or service in the middle of a pipeline?

Yes.

A DataZen pipeline can call HTTP/S endpoints while data is moving and use the response in subsequent processing.

That endpoint could be another SaaS application, an internal service, a cloud function, a machine-learning endpoint, or an AI service. The returned data can enhance, join with, or replace the current pipeline dataset.

Can applications push data directly into DataZen?

Yes.

DataZen supports incoming HTTP/S webhooks, allowing another application to send an event or payload directly into a pipeline.

The incoming data can then be transformed, captured, routed, and sent through downstream processing.

Can DataZen handle both batch and real-time integrations?

Yes. DataZen supports both batch and real-time integrations, with some limitations for the cloud edition.

DataZen supports a range of patterns, including full snapshots, high watermarks, Synthetic CDC, native CDC streams, overlapping windows, webhooks, messaging, and synchronization patterns.

This lets you choose the approach that best fits both the source system and how quickly the business actually needs the data.

AI and Agents

Where does AI fit into DataZen?

AI can operate directly inside a DataZen pipeline.

A pipeline can call an LLM or AI agent for classification, extraction, interpretation, enrichment, summarization, reasoning, or decision-making. The response can then immediately affect what happens next, such as applying filters and transformations, creating new fields, or triggering downstream actions—all inline within the same pipeline.

This means AI becomes part of the integration flow itself, rather than a separate process that only runs after the integration is finished.

You can send data to an AI Agent directly as part of the prompt or as a temporary file. For larger datasets, DataZen can write data to a temporary file, either in the cloud or locally, for the agent to read and then delete when the request is complete. This is known as a file-based claim-check pattern.

Is DataZen mainly an AI product?

No.

The underlying platform handles conventional integration, ETL/ELT, replication, CDC, APIs, databases, files, synchronization, and operational automation without requiring AI.

AI is an additional capability that can participate in those same pipelines when it adds value.

Do we have to use AI?

No.

A DataZen pipeline can be completely deterministic. You can use SQL, transformations, rules, APIs, database processing, CDC, and ordinary application logic without calling an AI model at all.

AI is another tool available to the pipeline, not a requirement.

Does using AI mean our data has to leave our environment?

Not necessarily.

DataZen calls AI through configured AI or HTTP endpoints rather than requiring one specific model provider. Where data goes, therefore, depends on the AI endpoint and deployment architecture your organization chooses.

That allows the AI portion of the architecture to be designed around your security, networking, privacy, and data-residency requirements.

Can DataZen turn an entire business process into a tool an AI agent can call?

Yes.

A DataZen Reader or Direct pipeline can be exposed as an MCP Tool with defined inputs, transformations, business logic, and outputs.

Instead of exposing the underlying systems directly to an agent, you can expose a governed business capability. For example, a pipeline might retrieve inventory and forecast data, reconcile them, and return the result through a single tool.

Put simply, the pipeline itself becomes the tool.

How is that different from just giving an AI agent a list of actions?

A list of actions gives an agent individual operations and leaves the agent responsible for deciding how to combine them.

DataZen lets you build the governed pipeline first, combining HTTP calls, database queries, SQL CDC transformations, parameters, and business logic. Once the pipeline is built, that complete, governed capability can be exposed to an AI agent as a single tool.

The agent interacts with the well-defined business operation rather than needing to understand every underlying system and step.

Can you limit AI Agents’ access to certain pipelines?

Yes.

AI Agents can, and should, be configured using Access Tokens generated by DataZen. Each token can then be associated with one or more topic areas. You can think of these topic areas as security groups that DataZen uses to control access. You can create as many topics as needed.

For example, you can create data pipelines that expose financial data under the “finance” topic, operational data under the “ops” topic, and general weather data under “general,” then create an access token that is limited to the “finance” and “general” topics.

Others create hundreds of MCP Tools automatically. Does DataZen do that?

No!

When it comes to business data, we believe less is more. Here is why.

Creating hundreds of MCP Tools for each individual system pushes a great deal of complexity onto the AI agent. The agent must then search through hundreds or even thousands of tools to determine which ones contain the right data and how to combine them. This increases the likelihood of an unhelpful result, and even when the agent gets the answer right, the process can consume a significant number of AI tokens and become expensive at scale.

These large tool catalogs make it much harder for AI Agents to determine the necessary sequence of steps and reproduce that same path from one request to the next.

Governance also becomes much harder to enforce when an AI Agent can access sources directly, bypassing the controls and boundaries necessary to keep data secure.

We believe your AI agents and users are better served when governed data pipelines serve as a secure data-servicing layer between your AI Agents and the underlying systems. Think of an MCP-enabled DataZen Pipeline as a controlled view into your business world. It can combine data from multiple systems when needed while limiting what data is exposed, enforcing governance, and monitoring every request.

The better approach is to create highly curated data pipelines and expose them as MCP Tools, making AI access more efficient, reliable, and secure.

Building and Operating Pipelines

Do I need to learn a new scripting language to use DataZen?

Not really. There is a small learning curve, but if you already know some SQL, DataZen will feel very familiar.

DataZen uses an integration language called SQL CDC. This is a simplified SQL-like scripting language designed specifically for routing data between systems, with advanced in-memory data transformation capabilities built in.

The syntax is based on familiar SQL concepts, with additional commands designed for integration tasks such as reading from sources, transforming data, capturing changes, calling external services, and pushing data to target systems.

Here is a simple example. This four-line SQL CDC pipeline retrieves weather alerts from an HTTP API, adds a new column on the fly in memory, and writes the results directly into a database. The JSON response is transformed into rows and columns as the pipeline runs. Records are appended automatically, and the target table is created if it does not already exist. All of this happens in four lines of SQL CDC code.

-- Get data from an HTTP API
SELECT * FROM HTTP [weather] (GET /alerts/active);

-- Transform it into rows and columns
APPLY TX 'features';

-- Add a column dynamically
ADD COLUMN 'string(25) State = @right({{Headline}}, 2)';

-- Append data to a table; the table is created if needed
SINK INTO DB [az-sql-db] TABLE 'weatherLatest' WITH APPEND;

DataZen provides a development guide, working examples, and an AI-accessible SQL CDC specification that compatible AI assistants can reference directly. This allows new users to start building business pipelines right away with AI-generated scripts while learning SQL CDC through practical use.

Deployment and Security

Can DataZen run in the cloud or inside our own environment?

Yes.

DataZen supports cloud-hosted agents as well as self-hosted agents for environments with different security, networking, and connectivity requirements.

Most core capabilities are available in both deployment models, although some connection types and features differ.

Can DataZen connect cloud systems with systems inside our network?

Yes.

DataZen supports architectures that span cloud and self-hosted environments. Self-hosted agents can operate close to internal applications or databases while other parts of the architecture interact with cloud services.

That allows organizations to connect systems across environments without first moving everything into the same place.

How does DataZen protect sensitive data?

DataZen applies security controls across agents, connections, APIs, and Change Logs.

Connections are stored in an encrypted format, service tokens can restrict programmatic access by scope, pipelines can mask sensitive fields, and X.509 certificates can be used for encryption, signing, and authentication.

These controls also let you apply different data protections depending on the destination—for example, masking sensitive fields before sending data into a test environment.