Skip to content

Data Actions

Everything Flowfile can do to your data is one of the actions on this page — read a file, filter rows, join two tables, summarize by group, train a model, write the result somewhere. You drag one onto the canvas, fill in its settings, and connect it to the next one. This page lists all of them, so you can see the full set without opening the app and find the one you need without knowing which category it lives in.

On the canvas, an action is called a node

The palette, the canvas and the rest of this reference use the word node. One node is one action, and the two words mean the same thing here.

Find an action by what you want to do

You want to Use Where
Keep only the rows that meet a condition Filter data Transformations
Add a calculated column Formula Transformations
Remove duplicate rows Drop duplicates Transformations
Fix blanks, stray spaces and inconsistent casing Data cleansing Transformations
Rename a lot of columns at once Rename columns Transformations
Choose, reorder or drop columns Select data Transformations
Number the rows Add record Id Transformations
Split one cell's list into separate rows Text to rows Transformations
Look up values from another table (VLOOKUP) Join Combine Operations
Stack two files that share the same columns Union data Combine Operations
Match names that are spelled slightly differently Fuzzy match Combine Operations
Run part of the flow only when a condition holds Gate Combine Operations
Reuse another flow as a single step Run Flow Combine Operations
Total or average per customer, month, region Group by Aggregations
Turn row values into columns (crosstab) Pivot data Aggregations
Turn columns back into rows Unpivot data Aggregations
Running total, rank or moving average Window functions Aggregations
Count the rows Count records Aggregations
Read every file in a folder List files Input Sources
Look at the data before deciding what to do Explore data Output Operations
Save a result colleagues can query and chart Write to Catalog Output Operations
Predict a number or a category Train Model, then Apply Model Machine Learning
Write SQL or Python instead of filling in a form SQL Query, Polars code, Python Script Transformations

Nothing here fits? Build your own node in the Node Designer, or install one someone else published.

The six categories

The palette groups actions the same way this reference does, under the same headings.

  • Input Sources


    10 actions. Get data in: files, folders, databases, cloud storage, REST APIs, Kafka, Google Analytics, the catalog.

  • Transformations


    14 actions. Reshape one dataset: filter, sort, cleanse, calculate columns, or drop into SQL, Polars or Python.

  • Combine Operations


    8 actions. Bring datasets together with a join, union or fuzzy match — group connected records, branch the flow with a gate, or call another flow.

  • Aggregations


    5 actions. Summarize and restructure: group, pivot, unpivot, count, window calculations.

  • Machine Learning


    4 actions. Split a dataset, fit a model, score new rows, and measure how well it did.

  • Output Operations


    7 actions. Send results out: files, databases, cloud storage, the catalog, an API response — or explore them on screen.

Every action, A to Z

48 actions as of 2026-09. The palette is the live list; this table is generated from the same source (flowfile_core/flowfile_core/configs/node_store/nodes.py) and each name matches what the palette shows.

Action What it does Category Lite
Add record Id Generate unique identifiers for each row Transformations ●
API response Return this dataset as the body of an HTTP API endpoint Output
Apply Model Score data with a trained model Machine Learning
Count records Calculate the total number of rows Aggregations ●
Cross join Create all possible combinations between two datasets Combine ●
Data cleansing Fix nulls, whitespace, unwanted characters and casing in one step Transformations
Drop duplicates Remove duplicate rows based on selected columns Transformations ●
Evaluate Model Compare actual vs predicted columns and compute quality metrics Machine Learning
Explore data Interactive data exploration and analysis Output ●
Filter data Keep only rows that match your conditions Transformations ●
Flow Input Named entry point for data when this flow runs inside another flow Input
Flow Output Named exit point exposing this dataset when the flow runs inside another flow Output
Formula Create or modify columns using custom expressions Transformations ●
Fuzzy match Join datasets based on similar values instead of exact matches Combine
Gate Pass data through only when a condition holds; otherwise skip what follows Combine
Google Analytics Load reports from a Google Analytics 4 property Input
Graph solver Group related records in graph-structured data Combine
Group by Aggregate data by grouping and calculating statistics Aggregations ●
Join Merge two datasets based on matching column values Combine ●
Kafka Source Read data from a Kafka or Redpanda topic Input
List files List a folder's contents as a table Input
Manual input Create data directly Input ●
Multi-field formula Apply one expression to many columns at once Transformations
Pivot data Convert data from long format to wide format Aggregations ●
Polars code Write custom Polars DataFrame transformations Transformations ●
Python Script Execute Python code on an isolated kernel container Transformations
Random Split Randomly partition rows into named groups (e.g. train/test) Machine Learning
Read data Load data from CSV, Excel, Parquet and other files Input ●
Read from Catalog Read a table from the data catalog Input ●
Read from cloud provider Read data from AWS S3 and other cloud storage Input
Read from Database Load data from database tables or queries Input
Rename columns Bulk-rename columns by prefix, suffix, or a formula Transformations ●
REST API Read JSON data from a REST API with auth and pagination Input
Run Flow Execute a flow from the catalog, mapping data and parameters into it Combine
Select data Choose, rename, and reorder columns to keep Transformations ●
Sort data Order your data by one or more columns Transformations ●
SQL Query Write SQL queries against connected data sources Transformations
Take Sample Work with a subset of your data Transformations ●
Text to rows Split text into multiple rows based on a delimiter Transformations
Train Model Fit a regression or classification model Machine Learning
Union data Stack multiple datasets by combining rows Combine ●
Unpivot data Transform data from wide format to long format Aggregations ●
Wait For Pass the left input through; the right input only enforces ordering Combine
Window functions Rolling, cumulative, rank, tile and partition-aggregate calculations Aggregations
Write data Save your data as CSV, Excel, Parquet and other files Output ●
Write to Catalog Save data as a table in the data catalog Output ●
Write to cloud provider Save data to AWS S3 and other cloud storage Output
Write to Database Save data to database tables Output

A ● marks the 22 actions that also run in Flowfile Lite, the browser-only edition. Lite adds two of its own — External Data and External Output, which fetch from and post to a URL — for 24 in total.

How an action works

Every node on the canvas behaves the same way:

  • Inputs and outputs. A node reads from whatever is connected to its input handles and passes its result on from its output handle. The count is fixed per action: Join takes two inputs, Random Split emits two outputs, Write data has no output at all.
  • Settings. Click a node and its settings open in the right-hand panel. The fields differ per action; each section below lists them.
  • Schema preview. Once configured, a node reports its output columns and types without running the flow. Run it to see actual rows in the preview panel.
  • Lazy by default. Most actions build up a Polars query that only executes when you run the flow, so intermediate steps cost nothing until you ask for a result.

New to the canvas? Building Flows covers creating, connecting, configuring and running nodes. Every action here is also available from Python.