Atlas / Skills / wgzhao / SKILL: Addax 项目知识

SKILL: Addax 项目知识SAFE

skills/wgzhao/addax

Actively maintained successor to Alibaba DataX — a fast, versatile, open-source ETL tool for 20+ RDBMS and NoSQL data sources.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
Apache-2.0
Stars
1,445
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

Addax

Any source. Any target. Fast.

Addax is an actively maintained successor to Alibaba's DataX, which has been frozen since 2023 — a fast, versatile, open-source ETL (Extract, Transform, Load) tool supporting over 30 RDBMS and NoSQL data sources. It provides a growing ecosystem of plugins and offers easy-to-follow configuration for data integrations.

简体中文

🚀 Features

  • Supports 30+ SQL and NoSQL data sources, and easily extendable for more.
  • Configurable via simple JSON-based job descriptions.
  • Actively maintained with improved architecture and added functionality compared to DataX.
  • Docker images for quick deployment.

📚 Documentation

Detailed instructions on installation, configuration, and usage are available:

💚 Project Status

Addax is actively maintained. The project is in a mature maintenance phase by design:

  • Monthly maintenance releases. A new release ships roughly every month (see Releases), bundling dependency/CVE updates and bug fixes — even when there are no new features to announce.
  • Responsive issue triage. Open issues are typical
Read from source at commit 240a3f9e0ca7OBSERVED · 2026-10-08
02

Install

Commands as the repository documents them. They are shown, not run.

git clone https://github.com/wgzhao/addax.git addax
git clone https://github.com/wgzhao/addax.git addax
git clone https://github.com/wgzhao/addax.git
03

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

# SKILL: Addax 项目知识

## 1. 项目整体认识

### 1.1 项目定位

- **名称**:Addax  
- **类型**:通用开源 ETL 工具(Extract–Transform–Load)  
- **起源**:基于阿里巴巴 DataX 的 fork 与演进  
- **目标**:在多种异构数据源之间,提供稳定、高效、可扩展的“离线数据同步”能力

### 1.2 核心价值

- 支持 **30+ SQL/NoSQL/文件/时序/大数据** 数据源
- 使用 **JSON 任务配置** 即可完成复杂同步,无需写代码
- 插件化架构,Reader / Writer / Transformer 解耦,可自由扩展
- 提供 **数据质量监控、速率控制、错误容忍、脏数据探测** 等生产级能力
- 既可命令行运行,也可通过 **Server 模块 HTTP 接口** 异步提交和管理任务
- 有配套的 **addax-admin / addax-ui** 项目做 Web 管控

---

## 2. 概念与架构模型

### 2.1 核心业务概念

在与用户讨论 / 理解需求时,应优先按以下抽象模型理解:

- **Job(作业)**
  - 一次完整的数据同步任务,从一个源到一个目标
  - 通过一个 JSON 文件描述:数据源 reader、目标端 writer、变换规则、速率控制、错误阈值等
  - Job 是业务上的最小单位,如 “从 MySQL 表 A 同步到 PostgreSQL 表 B”

- **Task(子任务)**
  - 为提升性能,将一个 Job 拆分为多个 Task 并发执行
  - 每个 Task 负责同步一部分数据(如若干分表、某一范围分片)

- **TaskGroup**
  - 一组 Task 的集合,由框架统一调度执行
  - 每个 TaskGroup 内有若干通道(channel),每个 channel 负责一条 `Reader → Channel → Writer` 流水线

- **Reader 插件**
  - 数据采集模块,负责从“源数据源”读取数据,发送给框架
  - 只关心“如何正确读”,不关注类型转换、指标统计等通用问题

- **Writer 插件**
  - 数据写入模块,负责从框架拿数据写入“目标端”
  - 只关心“如何正确写”,通用逻辑由框架处理

- **Transformer(数据转换)**
  - 可选模块,在 Reader 和 Writer 之间对数据进行转换
  - 支持内置 UDF:`dx_substr` / `dx_pad` / `dx_replace` / `dx_filter` / `dx_groovy`
  - 可以做脱敏、字段裁剪、补全、过滤、自定义 Groovy 脚本转换等

- **Channel(通道)**
  - Reader 到 Writer 之间的数据通路和缓冲队列
  - 决定并发度 & 流量控制(基于字节数、记录数、通道数)

### 2.2 架构概览

- **整体框架**:Framework + 插件(Reader / Writer / Transformer)
- 数据通路(简化):

  - **源端 → Reader → Framework(Channel) → Writer → 目标端**

- 作业生命周期(JobContainer 内部):
  1. `preHandler()` – 作业前置处理
  2. `init()` – 初始化 reader/writer 插件
  3. `prepare()` – 源端和目标端的准备工作
  4. `split()` – 按并发度拆分成多个 Task
  5. `schedule()` – 将 Task 组织为 TaskGroup,并发执行
  6. `post()` – 全局后置收尾(如 rename 影子表)
  7. `postHandler()` – 作业后置处理

- Task 执行:
  - 每个 Task 固定以 `Reader → Channel → Writer` 的线程模型执行
  - Channel 内以 Record/Column 为单位传输数据

---

## 3. SKILL:与用户交互时的“领域语言”

### 3.1 如何理解/解释一个 Job JSON

Job JSON 顶层结构:

```json
{
  "job": {
    "settings": {},
    "content": {
      "reader": {},
      "writer": {},
      "transformer": []
    }
  }
}
```

- `job.settings`
  - 控制本次任务的全局行为
  - 重点字段:
    - `speed.byte`:每秒允许的最大字节数(Bps),`-1` 表示不限制
    - `speed.record`:每秒允许的最大记录数
    - `speed.channel`:通道数(影响 Task 数量)
    - `errorLimit.record`:允许错误记录总数
    - `errorLimit.percentage`:允许错误记录占比

- `job.content.reader`
  - 必填,描述数据源及其读取方式
  - 核心字段(以关系型数据库为例):
    ```json
    {
      "name": "mysqlreader",
      "parameter": {
        "username": "",
        "password": "",
        "column": [],
        "autoPk": false,
        "splitPk": "",
        "connection": [
          {
            "jdbcUrl": [],
            "table": []
          }
        ],
        "where": ""
      }
    }
    ```

- `job.content.writer`
  - 必填,描述落地目标及写入策略
  - 核心字段(以关系库为例):
    ```json
    {
      "name": "mysqlwriter",
      "parameter": {
        "username": "",
        "password": "",
        "writeMode": "",
        "column": [],
        "session": [],
        "preSql": [],
        "postSql": [],
        "connection": [
          {
            "jdbcUrl": "",
            "table": []
          }
        ]
      }
    }
    ```

- `job.content.transformer`
  - 可选,列表形式,每个元素为一个转换规则
  - 典型片段(内置函数):
    ```json
    {
      "transformer": [
        {
          "name": "dx_substr",
          "parameter": { "idx": 1, "pos": 0, "length": 3 }
        }
      ]
    }
    ```

AI 在阅读/生成 Job 时,应显式区分:
- 全局控制(settings) vs 数据流定义(reader/writer) vs 转换规则(transformer)

### 3.2 任务拆分与并发推导逻辑

- 用户设定并发度:`job.settings.speed.channel = N`
- 框架内部:
  1. Reader 的 `split()` 按源端特性拆成若干 Task(如按分表、分片、主键范围等)
  2. Writer 的 `split()` 需与 Reader 的 Task 数量 **1:1 对齐**
  3. Scheduler 根据 `taskGroup.channel`(`conf/core.json` 中配置)决定 TaskGroup 数量:
     - `taskGroupCount = speed.channel / taskGroup.channel`
- 与用户讨论“为什么任务这么慢/这么多连接”时,应从:
  - `speed.channel`
  - Reader 拆分策略(是否按 `splitPk` / 分区表)
  - 目标端写入瓶颈(Writer 能力、批量大小等)
  入手解释。

### 3.3 数据质量与错误处理

Addax 在数据质量方面的关键点:

- **类型不丢失/不失真**
  - 内部抽象了统一的 Column 类型:`Long / Double / String / Date / Timestamp / Bool / Bytes`
  - 每个插件有自己的类型转换策略,保证最小损失

- **错误控制**
  - 通过 `errorLimit.record` 和 `errorLimit.percentage` 控制“可容忍错误”
  - 超过阈值即认为任务失败

- **脏数据(Dirty Data)**
  - 概念:传输过程中因各种原因(例如类型不匹配)导致出错的记录
  - 能够过滤、识别、收集与展示脏数据,并统计数量和字节数
  - Transformer 层若抛出异常/返回 null,也会影响成功/失败/过滤计数

AI 在帮助用户排错时:
- 应主动询问/检查:错误是否集中在类型转换、特定列、特定插件
- 建议合理设置 `errorLimit`,在保障数据质量和任务稳定之间平衡

---

## 4. 使用方式与运行环境

### 4.1 安装与运行

- **运行时环境**
  - Java:JDK 17
  - Python 2.7+ / 3.7+(仅 Windows 使用本地脚本时需要)

- **三种典型使用方式**

1)Docker 运行示例:
```bash
docker pull quay.io/wgzhao/addax:latest
docker run -ti --rm --name addax \
  quay.io/wgzhao/addax:latest \
  /opt/addax/bin/addax.sh /opt/addax/job/job.json
```

2)一键安装脚本(Linux / macOS):
```bash
/bin/bash -c "$(curl -fsSL https://raw.githubusercontent.com/wgzhao/Addax/master/install.sh)"
```
- 安装目录:
  - macOS:`/usr/local/addax`
  - Linux:`/opt/addax`

3)源码编译:
```bash
git clone https://github.com/wgzhao/addax.git
cd addax
mvn clean package
mvn package -Pdistribution        # 或 assembly:single
# 产出目录示例:target/addax-<version>
```

- **首次运行任务**
  ```bash
  bin/addax.sh job/job.json
  ```

### 4.2 命令行工具 addax.sh

基本用法:
```bash
bin/addax.sh <job_file> [options]
```

关键参数(AI 在给终端命令建议时要正确使用):

- `-h, --help`:帮助
- `-v, --version`:版本
- `-l, --log`:指定日志文件路径
- `-d, --debug`:开启调试模式(IDEA 远程调试会用到)
- `-L, --log-level`:`DEBUG | INFO | WARN | ERROR`
- `-j, --jvm`:追加 JVM 参数
- `-p, --params`:向 Job 传入动态参数(`-Dkey=value` 形式)

示例(动态参数):
```bash
bin/addax.sh job/test.json \
  -p "-Dusername=root -Dpassword=123456 -Dparam1=value1 -Dparam2=value2"
```

Job JSON 中可以通过 `${param}` 访问这些值,比如 `${username}`。

内置时间变量(示例时间 `2025-07-16 12:13:14`):
- `${curr_date_short}` → `20250716`
- `${curr_date_dash}` → `2025-07-16`
- `${curr_datetime_short}` → `20250716121314`
- `${curr_datetime_dash}` → `2025-07-16 12:13:14`
- `${biz_date_short}` / `${biz_date_dash}` / `${biz_datetime_*}` 等

### 4.3 Server 模块(HTTP 提交任务)

- 用途:通过 HTTP 提交 Job JSON 并异步执行,可查询进度和结果
- 启动脚本:`core/src/main/bin/addax-server.sh`

启动示例:
```bash
./addax-server.sh start        
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeWARN
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
declared (2 observation(s))
Network
declared (5 observation(s))
Shell
declared (1 observation(s))
Dependencies
pinned
Secrets in source
none-found

Findings (12)

MEDIUMObfuscation / stealth · obf.base64_blob · CWE-506, CWE-94
images/addax-flowchart.drawio:1
<mxfile host="Electron" modified="2023-02-05T08:34:16.961Z" agent="5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) draw.io/20.8.16 Chrome/106.0.5249.199 Electron/21.4.0
LOWInventory / provenance · inv.symlink · CWE-1104
CLAUDE.md
CLAUDE.md
Why it matters. link not followed
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
shrink_package.sh:43
( cd ${plugin_dir} && ln -sf ../../../../shared/${file_name} $file_name )
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
e2e/README.md:26
`E2E_ES_ENDPOINT` (`http://127.0.0.1:9200` by default). There is no throwaway cluster in
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
e2e/cases/250_s3reader_objects/setup.sh:65
s3 = boto3.client("s3", endpoint_url="http://127.0.0.1:%s" % port,
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
e2e/cases/270_elasticsearchreader_scroll/setup.sh:9
ES_ENDPOINT="${E2E_ES_ENDPOINT:-http://127.0.0.1:9200}"
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
e2e/cases/280_elasticsearchwriter_types/setup.sh:9
ES_ENDPOINT="${E2E_ES_ENDPOINT:-http://127.0.0.1:9200}"
LOWNetwork egress · net.raw_ip · CWE-200, CWE-319
e2e/cases/280_elasticsearchwriter_types/verify.sh:26
ES_ENDPOINT="${E2E_ES_ENDPOINT:-http://127.0.0.1:9200}"
INFOInventory / provenance · inv.oversize · CWE-1104
plugin/reader/tdenginereader/src/main/libs/libtaos.so.2.0.16.0
plugin/reader/tdenginereader/src/main/libs/libtaos.so.2.0.16.0
Why it matters. 8085768 bytes not read
INFOInventory / provenance · inv.oversize · CWE-1104
plugin/writer/tdenginewriter/src/main/libs/libtaos.so.2.0.16.0
plugin/writer/tdenginewriter/src/main/libs/libtaos.so.2.0.16.0
Why it matters. 8085768 bytes not read
INFOSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
release-notes/6.1.0.md:60
- `install.sh` reads prompts from the terminal, so `curl -fsSL ... | bash` works (#1554); deploys create missing remote directories (#1553); `build-module.sh` base paths and CLI options are repaired.
INFOSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
release-notes/6.1.0.md:143
- `install.sh` 从终端读取交互输入,`curl -fsSL ... | bash` 可以正常使用 (#1554);部署时会创建缺失的远端目录 (#1553);修复 `build-module.sh` 的基础路径与命令行选项。

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 240a3f9e0ca7full audit observations/trust-audit/skill/wgzhao__addax.json · Report an issue / request a re-scan
05

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-08240a3f9e0ca7SAFEB89first audit
06

Questions

What does the SKILL: Addax 项目知识 skill do?

Actively maintained successor to Alibaba DataX — a fast, versatile, open-source ETL tool for 20+ RDBMS and NoSQL data sources.

Is SKILL: Addax 项目知识 safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can SKILL: Addax 项目知识 access on my machine?

The audit observed that it reaches the network, runs shell commands and reads or writes files. Each of those is consistent with what it says it does. Secrets in the source: none found.

How current is this page?

The grade is for one exact copy of the source (240a3f9e0ca7), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement