All Interzoid products and tools from a single launch point: Quickly solve data challenges - better ROI for everything your data flows into -> Launch Now!

High-Performance Databricks Batch API Processing

Match organization names, individual names, addresses, and general text across entire lakehouse tables

Lightning-fast parallel processing with AI-powered similarity matching, directly against your SQL warehouse

Built for Databricks

Transform Your Data Processing Workflow

Leverage our high-performance, parallel-processing cloud architecture to run AI-driven matching jobs against entire lakehouse tables or views, finding the duplicate and near-duplicate records that exact matching cannot.

10x
Faster, Parallel Processing
High
Volume Tables Processed at Scale
Zero
Data Exports or Copies Required

Works With Every Databricks Deployment

Interzoid connects through a SQL warehouse using a personal access token, so it reads from any workspace on any cloud, whether governed by Unity Catalog or a legacy metastore.

Databricks on AWS
Azure Databricks
Databricks on Google Cloud
Unity Catalog Catalog, schema, and table
hive_metastore Legacy workspaces supported
Serverless SQL Warehouses
Classic and Pro Warehouses
Delta Lake Tables
Views and Materialized Views
Delta Sharing Recipients Shared data read in place
Any Workspace Region
HTTPS on Port 443 No firewall rule needed

No connector to install, no notebook to run, and nothing deployed inside your workspace. Supply a workspace hostname, a SQL warehouse HTTP path, and an access token, and Interzoid reads only the columns you select. Because the connection is HTTPS on port 443, no special outbound firewall rule is required.

Powerful Features for Modern Data Teams

Lakehouse Table Integration

Read Delta Lake tables and views directly through a SQL warehouse, governed by Unity Catalog permissions, with a dedicated read-only service principal recommended.

Learn More

AI-Powered Data Matching

Advanced algorithms and AI models identify similar data elements, generate similarity keys for clustering matches and detecting inconsistencies within and across your data tables.

Learn More

Parallel Processing Engine

Multi-threaded, high-performance architecture processes high volumes of data in seconds using our distributed cloud infrastructure.

Learn More

How It Works

Get started in minutes with our streamlined three-step process

1

Connect to Databricks

Supply your workspace hostname, SQL warehouse HTTP path, and access token, or paste a connection string, then pick a catalog and schema.

2

Select Your Matching Algorithms

Choose the algorithm that fits your data: organization names, individual names, addresses, general text, or a combination of two of them at once.

3

Download Results

Receive a match report with records grouped into clusters by similarity key, ready to review, share, or load back into your lakehouse tables.

Data Matching Example

See our Databricks batch processing in action - generating similarity keys to match company/organization name data from directly within lakehouse tables.

Interzoid match report from a Databricks table showing company name variations grouped into clusters by similarity key

What You Can Match

Four data types, plus combinations that raise precision when one field alone is too broad

Organization Names

Resolve Acme Corp, ACME Corporation, and Acme Inc. to a single organization, across the legal suffixes, punctuation, abbreviations, and spelling variations that make company names so inconsistent.

Individual Names

Match James Johnston, Jim Johnston, and J. Johnston as the same person, handling nicknames, initials, middle names, and name ordering.

Street Addresses

Reconcile 400 E Broadway St with 400 East Broadway Street, resolving directional abbreviations, street type variations, unit designations, and spacing.

General Text

Generate similarity keys for other text values, so product names, descriptions, and free-form fields cluster the same way that names and addresses do.

Combination Matching

Require two fields to agree at once: organization with address, organization with individual name, or address with individual name. This separates the Dallas branch from the Phoenix branch while still recognizing that both spell the parent company four different ways.

Similarity Keys and Clusters

Every value receives a key representing the entity behind the text rather than the characters themselves. Records sharing a key are grouped into clusters, which makes each match auditable instead of something to take on faith.

Frequently Asked Questions

Which Databricks deployments are supported?

Databricks workspaces on AWS, Azure, and Google Cloud, in any region, reading through a serverless, pro, or classic SQL warehouse. Unity Catalog workspaces and legacy hive_metastore workspaces are both supported, as are Delta Lake tables, views, materialized views, and data shared in through Delta Sharing.

Does Interzoid modify my database?

No. Processing issues SELECT statements against the table and columns you choose and nothing else. Nothing is written, updated, or deleted, no tables are created, and no job or notebook runs in your workspace. Results are returned to your browser, and you decide what to do with them.

Do I need to export my data first?

No. Interzoid connects directly to your SQL warehouse and reads the columns you select. There is no CSV export, no staging table, no cloud storage copy, and nothing to secure, track, or delete afterward.

How is my access token handled?

The personal access token is transmitted over HTTPS with each request and is not retained after the session ends. Because a token inherits the permissions of whoever created it, and because processing is read-only, the recommended practice is a dedicated service principal granted only USE CATALOG, USE SCHEMA, and SELECT, which can be revoked at any time.

How is usage measured?

Each record processed consumes one Interzoid API credit. Trial credits are included with a new API account, so a first table can be processed at no cost, and high volume tables are supported for production workloads.

Can I run this from my own applications instead of the browser?

Yes. The same processing is available as a REST API that can be called from scripts, applications, and data pipelines, which allows Databricks matching jobs to be automated without using the browser wizard.

Learn More

Documentation, background reading, and the rest of the Interzoid platform

Ready to Transform Your Data Quality?

Join hundreds of data teams already using Interzoid's batch processing APIs

Start Processing Now Read Documentation

Questions? Contact our team at support@interzoid.com