Quickstart¶
Status: Available in ETLantic 0.53.0 (Beta release candidate). Use
python -m etlantic initfor the recommended CLI-first path with durable reports and declarative assets. Budget ~5–10 minutes for first success; optional validation aha below adds a few minutes.
PyPI vs clone
This page is for PyPI installs. Repository examples/ scripts need a
git checkout and uv — see Installation.
1. Install¶
ETLantic requires Python 3.11 or newer. Prefer python -m so PATH issues do not
block you. See Installation for full options.
Unix / macOS:
python -m venv .venv && source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install 'etlantic==0.53.0'
python -m etlantic --version # expect 0.53.0
Windows (PowerShell):
py -3.11 -m venv .venv
.\.venv\Scripts\Activate.ps1
py -3.11 -m pip install --upgrade pip
py -3.11 -m pip install 'etlantic==0.53.0'
py -3.11 -m etlantic --version
2. Initialize a project¶
init requires an empty directory (or pass --force if the directory
already has files, for example a Poetry or uv project):
--with-toml writes a minimal etlantic.toml (project name + default profile)
so later CLI invocations resolve project settings without extra flags. You can
omit it for a profile-JSON-only scaffold.
This creates pipeline.py (SamplePipeline), profiles/development.json,
sample data/sample.json, and .etlantic/ workspace directories.
The generated profile uses dataframe_engine: "local" — the built-in local
Python runtime (not Polars/Pandas) — and
portable_transform_policy: "require". The generated transformation is an
engine-neutral ETLantic definition, so changing engines later does not require
rewriting its body. Add an engine via Engine selection.
3. Validate and run (first success)¶
Pipeline targets use path/to/file.py:PipelineClass,
package.module:PipelineClass, or a path to an etlantic.pipeline/1 JSON
document. See CLI — Pipeline targets.
python -m etlantic validate pipeline.py:SamplePipeline --profile development
python -m etlantic run pipeline.py:SamplePipeline --profile development
cat data/out.json
No Python-side seed is required: the generated profile maps assets to
json://data/... paths.
What success looks like¶
validateexits 0 and prints a report with no errors (add--format jsonfor machines). Typical text output includes a summary line with 0 errors.runprints a run status ofsucceeded(look forstatus: succeededor equivalent in the run summary).data/out.jsoncontains Ada and Grace (identity transform on the sample).
Expected shape:
Optional later: python -m etlantic doctor --profile development,
inspect, plan, and report list.
4. Optional — see validation catch a bug¶
Skip if you only want the five-minute green path; return here when you want to feel validate-before-write.
The etlantic init scaffold defines Identity in pipeline.py (it is not
imported from etlantic). Its @Identity.portable body is compiled by the
selected engine. Edit only the Load annotation so the load expects a
different contract than the upstream step produces.
Before (generated):
import etlantic as etl
class Row(etl.Data):
id: int
name: str
class Identity(etl.Transformation):
rows: etl.Input[Row]
result: etl.Output[Row]
@Identity.portable
def identity(rows):
return rows
class SamplePipeline(etl.Pipeline):
raw: etl.Extract[Row] = etl.Extract(asset="rows")
step = Identity.step(rows=raw)
out: etl.Load[Row] = etl.Load(input=step.result, asset="out")
After (broken on purpose — add Other and change only the Load line):
import etlantic as etl
class Row(etl.Data):
id: int
name: str
class Other(etl.Data):
id: int
name: str
class Identity(etl.Transformation):
rows: etl.Input[Row]
result: etl.Output[Row]
@Identity.portable
def identity(rows):
return rows
class SamplePipeline(etl.Pipeline):
raw: etl.Extract[Row] = etl.Extract(asset="rows")
step = Identity.step(rows=raw)
# Broken: Load expects Other but step.result is still Row
out: etl.Load[Other] = etl.Load(input=step.result, asset="out")
Optional equivalent as a unified diff (same scaffold imports):
class Row(Data):
id: int
name: str
+class Other(Data):
+ id: int
+ name: str
+
+
class Identity(Transformation):
rows: Input[Row]
result: Output[Row]
@@
step = Identity.step(rows=raw)
- out: Load[Row] = Load(input=step.result, asset="out")
+ out: Load[Other] = Load(input=step.result, asset="out")
Re-validate:
Expect a non-zero exit and a wiring diagnostic such as:
data/out.json must not gain a new successful write until you restore
Load[Row]. That is the product promise: validate before write.
Restore out: etl.Load[Row] = etl.Load(input=step.result, asset="out") (and remove
Other if unused). Continue with an intentional uppercase transform in
First Pipeline—you can skip the wiring demo there if you
just completed this step.
5. Python SDK path (optional)¶
From the same project directory created by init (so pipeline.py is
importable as a top-level module):
from pipeline import SamplePipeline
report = SamplePipeline.validate(profile="development")
report.raise_for_errors()
SamplePipeline.run(profile="development")
If you see ModuleNotFoundError: pipeline, cd into the init project root
(the directory that contains pipeline.py) and retry.
Standards acronyms (ODCS / DTCS / DPCS) and Gate A/B labels appear later in Capabilities and Foundations—you do not need them for first success.