Turn enterprise data into a foundational asset for the AI era

Simba Data Foundation

Simba Data Foundation is one of two core pillars of the enterprise AGI platform, alongside the model foundation. It connects enterprise-scale data assets with AI capabilities, bringing together data ingestion, governance, knowledge capture, and secure delivery so models and intelligent business applications can access trusted, governed data on demand.

Let data flow in ways that models can understand, businesses can reuse, and security teams can audit.

The data foundation goes beyond a traditional data platform. It organizes enterprise data into production-ready assets that AI and other IT systems can reliably consume and reuse.

01

Data Built for AI

Redefine how data is organized, accessed, and protected so enterprise data is ready for foundation model use cases.

02

Data Stays in Your Domain

Support full-stack private deployment so enterprise data remains within your environment, helping protect data sovereignty and support compliance.

03

Knowledge as an Asset

Turn metrics, business rules, industry practices, and expert knowledge into structured assets AI can reuse.

04

Connect Once, Reuse Everywhere

Unify integration, governance, and delivery to reduce duplicate work and accelerate the journey from AI pilots to scaled deployment.

Store More. Compute Faster. Run Reliably. Govern Clearly. Control Costs.

Keep high-value data ready to support models and business decisions.

500B records/day

Data Throughput

Deliver stable, model-facing throughput for EB- and PB-scale data flows.

1M+ QPS

Query Throughput

Low-latency, high-throughput data APIs through a unified interface, supporting 100K+ QPS.

100B+ events

Real-Time Processing

Unified stream and batch processing, delivering query results within seconds—even for complex queries spanning large datasets, multiple sources, joins, and business rules.

1/3-1/5

Storage Compression

High-efficiency encoding reduces storage requirements to one-third to one-fifth of their original size.

Simba Data Foundation: A Governed Data Hub for AI Models

Bring together enterprise data sources, governance systems, knowledge assets, and access controls to provide AI applications with governed, trusted, on-demand data context.

AI Applications / Agents / Business Scenarios

Unified access
DeciAGIDataAGIOpsAGIOmnichannel AllocationStore IntelligenceSales ForecastProduct RecommendationSupply Chain IntelligenceBusiness DecisionsEnterprise Knowledge Assistant

Unified Data Service Interface

Low Latency, High Throughput, and Auditability
Model data APIsContext buildingOn-Demand Schema AccessCaching and accelerationPermission checksAudit trails
Multi-Source Data Access

Unify databases, warehouses, data lakes, APIs, files, and streaming data.

Unified Data Management

Identify schemas, field semantics, and relationships while managing catalogs, lineage, and lifecycle.

Data FoundationMetadata Center · Semantic Layer · Metrics System
Enterprise Knowledge Assets

Transform metric definitions, industry rules, business logic, and expert knowledge into structured, reusable AI assets.

Security and Access Governance

Bring together row- and column-level permissions, on-demand context loading, audit trails, and private deployment controls.

Enterprise Multi-Source Data

Unified ingestion · Unified governance
Relational databasesData warehousesData lakesBusiness APIsDocumentsStreaming data
InIngestion
Connect sources in minutes
GovGovernance
Manage metadata, lineage, and quality
KnKnowledge
Reuse metrics, rules, and expertise
SecSecurity
Private deployment, permissions, and audits
APIDelivery
Reliable model-facing services
Positioning

Traditional data platforms serve people. This data foundation serves models. It helps models understand business terms, honor access controls, access trusted data, and continuously turn enterprise knowledge into reusable intelligence assets.

Build a secure, trusted data flow from ingestion to model calls.

Simba brings together schemas, metadata, metrics, semantic layers, access policies, and caching to create a model-ready data service layer, so downstream applications receive stable, accurate, and auditable data context.

Multi-Source Data Access

Connect relational databases, data warehouses, data lakes, APIs, files, and streaming data using low-code connectors.

Unified Governance and Knowledge Capture

Ingest schemas, manage metadata and lineage, and capture metrics, semantic layers, industry rules, and expert practices.

Secure Delivery to Models and Agents

Use row- and column-level permissions, on-demand context loading, and audit trails so models only see authorized, business-aligned data.

Simba Data Foundation helps AI understand your business, use real data, and respect defined boundaries—not just answer questions.

General-purpose models are widely available, but AI that understands your business, uses your data, and respects your boundaries must be built on your own enterprise data foundation. Simba turns data assets into a continuously evolving enterprise intelligence layer.

DataBlack

DataBlack: Full-Lifecycle Data Security Engine

DataBlack is StartDT's data security engine. Built on a data-centric security architecture, it helps enterprises protect data across the full lifecycle, strengthen governance and risk management, and safeguard data assets.

Core Technologies

More comprehensive. More intelligent.

Full-Lifecycle Control

Apply comprehensive security controls across data flows and sensitive data.

Intelligent Classification

Apply intelligent algorithms and trained models to identify sensitive data faster and more accurately.

Comprehensive Audit

Provide detailed audit records, including sensitive-data usage, to support security and compliance requirements.

Masking and Encryption

Support static and dynamic masking, encryption, and authorized decryption to protect sensitive data in use, at rest, and in transit.

Risk Detection and Monitoring

Monitor high-risk user actions with custom rules or intelligent algorithms, and trace activity through security audit logs.

Product Architecture

Security Throughout the Data Lifecycle

DataBlack architecture: before-use identification includes sensitive-data discovery, masking, and encryption; in-use control covers dynamic control and risk control; post-use audit covers monitoring reports, sensitive-data usage records, and abnormal operation reports. Permission management runs through the full process, covering data permission policies, role-based permissions, permission change records, and project isolation.
DataKun

DataKun: Independently Managed Data Storage and Compute Engine

DataKun is an independently managed storage and compute engine that helps enterprises build a lightweight, intelligent data platform and develop their own analytics and processing capabilities. It supports multiple big-data jobs and services, custom components, validated version combinations, and continuous upgrades of core components.

Product Advantages

Technology You Control. Cost Efficiency.

Ready to Use Out of the Box

  • Deploy within an hour and create clusters in minutes.

One-Click Migration

  • Compatible with major open-source big data ecosystems.
  • Cluster migration tools help customers migrate quickly and smoothly.

Controlled Technology

  • Storage layer: supports HDFS and StartDT's SFS, which separates compute from storage.
  • Engine layer: supports community-supported open-source and custom components with ongoing updates.

Intelligent Operations

  • Layered operations and rich metrics help teams quickly detect and diagnose issues through visual monitoring.

Cost Efficiency

  • Separating compute from storage reduces storage costs.
  • Operations management tools reduce operating costs and manual effort.
  • Lower license costs reduce vendor lock-in risk.

Secure and Reliable

  • Supports high availability and high performance.
  • Enterprise-grade data security capabilities support security, reliability, and compliance requirements.

Product Architecture

Open, manageable, extensible, and continuously evolving

DataKun architecture: connected to DataSimba, with multi-engine compute systems such as Hive, Flink, Spark, Impala, ClickHouse, Presto, Kylin, YARN, and Phoenix; distributed storage such as SFS, HDFS, OSS, and OBS; data security with LDAP, Kerberos, and Ranger; and intelligent operations with cluster management, monitoring alerts, intelligent inspection, and diagnostics. Supported environments include Alibaba Cloud, Huawei Cloud, Tencent Cloud, JD Cloud, China Telecom Cloud, AWS, Azure, and on-premises IDC.

Data Foundation: making enterprise data governed, trusted, and available on demand for every AI scenario.