Data Actions
Everything Flowfile can do to your data is one of the actions on this page — read a file, filter rows, join two tables, summarize by group, train a model, write the result somewhere. You drag one onto the canvas, fill in its settings, and connect it to the next one. This page lists all of them, so you can see the full set without opening the app and find the one you need without knowing which category it lives in.
On the canvas, an action is called a node
The palette, the canvas and the rest of this reference use the word node. One node is one action, and the two words mean the same thing here.
Find an action by what you want to do
| You want to | Use | Where |
|---|---|---|
| Keep only the rows that meet a condition | Filter data | Transformations |
| Add a calculated column | Formula | Transformations |
| Remove duplicate rows | Drop duplicates | Transformations |
| Fix blanks, stray spaces and inconsistent casing | Data cleansing | Transformations |
| Rename a lot of columns at once | Rename columns | Transformations |
| Choose, reorder or drop columns | Select data | Transformations |
| Number the rows | Add record Id | Transformations |
| Split one cell's list into separate rows | Text to rows | Transformations |
| Look up values from another table (VLOOKUP) | Join | Combine Operations |
| Stack two files that share the same columns | Union data | Combine Operations |
| Match names that are spelled slightly differently | Fuzzy match | Combine Operations |
| Run part of the flow only when a condition holds | Gate | Combine Operations |
| Reuse another flow as a single step | Run Flow | Combine Operations |
| Total or average per customer, month, region | Group by | Aggregations |
| Turn row values into columns (crosstab) | Pivot data | Aggregations |
| Turn columns back into rows | Unpivot data | Aggregations |
| Running total, rank or moving average | Window functions | Aggregations |
| Count the rows | Count records | Aggregations |
| Read every file in a folder | List files | Input Sources |
| Look at the data before deciding what to do | Explore data | Output Operations |
| Save a result colleagues can query and chart | Write to Catalog | Output Operations |
| Predict a number or a category | Train Model, then Apply Model | Machine Learning |
| Write SQL or Python instead of filling in a form | SQL Query, Polars code, Python Script | Transformations |
Nothing here fits? Build your own node in the Node Designer, or install one someone else published.
The six categories
The palette groups actions the same way this reference does, under the same headings.
-
10 actions. Get data in: files, folders, databases, cloud storage, REST APIs, Kafka, Google Analytics, the catalog.
-
14 actions. Reshape one dataset: filter, sort, cleanse, calculate columns, or drop into SQL, Polars or Python.
-
8 actions. Bring datasets together with a join, union or fuzzy match — group connected records, branch the flow with a gate, or call another flow.
-
5 actions. Summarize and restructure: group, pivot, unpivot, count, window calculations.
-
4 actions. Split a dataset, fit a model, score new rows, and measure how well it did.
-
7 actions. Send results out: files, databases, cloud storage, the catalog, an API response — or explore them on screen.
Every action, A to Z
48 actions as of 2026-09. The palette is the live list; this table is generated from the same source (flowfile_core/flowfile_core/configs/node_store/nodes.py) and each name matches what the palette shows.
| Action | What it does | Category | Lite |
|---|---|---|---|
| Add record Id | Generate unique identifiers for each row | Transformations | ● |
| API response | Return this dataset as the body of an HTTP API endpoint | Output | |
| Apply Model | Score data with a trained model | Machine Learning | |
| Count records | Calculate the total number of rows | Aggregations | ● |
| Cross join | Create all possible combinations between two datasets | Combine | ● |
| Data cleansing | Fix nulls, whitespace, unwanted characters and casing in one step | Transformations | |
| Drop duplicates | Remove duplicate rows based on selected columns | Transformations | ● |
| Evaluate Model | Compare actual vs predicted columns and compute quality metrics | Machine Learning | |
| Explore data | Interactive data exploration and analysis | Output | ● |
| Filter data | Keep only rows that match your conditions | Transformations | ● |
| Flow Input | Named entry point for data when this flow runs inside another flow | Input | |
| Flow Output | Named exit point exposing this dataset when the flow runs inside another flow | Output | |
| Formula | Create or modify columns using custom expressions | Transformations | ● |
| Fuzzy match | Join datasets based on similar values instead of exact matches | Combine | |
| Gate | Pass data through only when a condition holds; otherwise skip what follows | Combine | |
| Google Analytics | Load reports from a Google Analytics 4 property | Input | |
| Graph solver | Group related records in graph-structured data | Combine | |
| Group by | Aggregate data by grouping and calculating statistics | Aggregations | ● |
| Join | Merge two datasets based on matching column values | Combine | ● |
| Kafka Source | Read data from a Kafka or Redpanda topic | Input | |
| List files | List a folder's contents as a table | Input | |
| Manual input | Create data directly | Input | ● |
| Multi-field formula | Apply one expression to many columns at once | Transformations | |
| Pivot data | Convert data from long format to wide format | Aggregations | ● |
| Polars code | Write custom Polars DataFrame transformations | Transformations | ● |
| Python Script | Execute Python code on an isolated kernel container | Transformations | |
| Random Split | Randomly partition rows into named groups (e.g. train/test) | Machine Learning | |
| Read data | Load data from CSV, Excel, Parquet and other files | Input | ● |
| Read from Catalog | Read a table from the data catalog | Input | ● |
| Read from cloud provider | Read data from AWS S3 and other cloud storage | Input | |
| Read from Database | Load data from database tables or queries | Input | |
| Rename columns | Bulk-rename columns by prefix, suffix, or a formula | Transformations | ● |
| REST API | Read JSON data from a REST API with auth and pagination | Input | |
| Run Flow | Execute a flow from the catalog, mapping data and parameters into it | Combine | |
| Select data | Choose, rename, and reorder columns to keep | Transformations | ● |
| Sort data | Order your data by one or more columns | Transformations | ● |
| SQL Query | Write SQL queries against connected data sources | Transformations | |
| Take Sample | Work with a subset of your data | Transformations | ● |
| Text to rows | Split text into multiple rows based on a delimiter | Transformations | |
| Train Model | Fit a regression or classification model | Machine Learning | |
| Union data | Stack multiple datasets by combining rows | Combine | ● |
| Unpivot data | Transform data from wide format to long format | Aggregations | ● |
| Wait For | Pass the left input through; the right input only enforces ordering | Combine | |
| Window functions | Rolling, cumulative, rank, tile and partition-aggregate calculations | Aggregations | |
| Write data | Save your data as CSV, Excel, Parquet and other files | Output | ● |
| Write to Catalog | Save data as a table in the data catalog | Output | ● |
| Write to cloud provider | Save data to AWS S3 and other cloud storage | Output | |
| Write to Database | Save data to database tables | Output |
A ● marks the 22 actions that also run in Flowfile Lite, the browser-only edition. Lite adds two of its own — External Data and External Output, which fetch from and post to a URL — for 24 in total.
How an action works
Every node on the canvas behaves the same way:
- Inputs and outputs. A node reads from whatever is connected to its input handles and passes its result on from its output handle. The count is fixed per action: Join takes two inputs, Random Split emits two outputs, Write data has no output at all.
- Settings. Click a node and its settings open in the right-hand panel. The fields differ per action; each section below lists them.
- Schema preview. Once configured, a node reports its output columns and types without running the flow. Run it to see actual rows in the preview panel.
- Lazy by default. Most actions build up a Polars query that only executes when you run the flow, so intermediate steps cost nothing until you ask for a result.
New to the canvas? Building Flows covers creating, connecting, configuring and running nodes. Every action here is also available from Python.