Flowfile Core API Reference
This section provides a detailed API reference for the core Python objects, data models, and API routes in flowfile-core. The documentation is generated directly from the source code docstrings.
Core Components
This section covers the fundamental classes that manage the state and execution of data pipelines. These are the main "verbs" of the library.
FlowGraph
The FlowGraph is the central object that orchestrates the execution of data transformations. It is built incrementally as you chain operations. This DAG (Directed Acyclic Graph) represents the entire pipeline.
flowfile_core.flowfile.flow_graph.FlowGraph
A class representing a Directed Acyclic Graph (DAG) for data processing pipelines.
It manages nodes, connections, and the execution of the entire flow.
Methods:
| Name | Description |
|---|---|
__init__ |
Initializes a new FlowGraph instance. |
__repr__ |
Provides the official string representation of the FlowGraph instance. |
add_api_response |
Adds an API-response sink node. |
add_apply_model |
Adds an Apply Model node. |
add_catalog_reader |
Adds a node that reads a table from the catalog. |
add_catalog_writer |
Adds a node that writes its input to the catalog as a Delta table or virtual table. |
add_cloud_storage_reader |
Adds a cloud storage read node to the flow graph. |
add_cloud_storage_writer |
Adds a node to write data to a cloud storage provider. |
add_cross_join |
Adds a cross join node to the graph. |
add_database_reader |
Adds a node to read data from a database. |
add_database_writer |
Adds a node to write data to a database. |
add_datasource |
Adds a data source node to the graph. |
add_dependency_on_polars_lazy_frame |
Adds a special node that directly injects a Polars LazyFrame into the graph. |
add_dynamic_rename |
Adds a node that renames many columns at once via a single rule. |
add_evaluate_model |
Adds an Evaluate Model node. |
add_explore_data |
Adds a specialized node for data exploration and visualization. |
add_external_source |
Adds a node for a custom external data source. |
add_filter |
Adds a filter node to the graph. |
add_flow_input |
Adds a named subflow-input placeholder source. |
add_flow_output |
Adds a named subflow-output sink (passthrough, always materialized). |
add_formula |
Adds a node that applies a formula to create or modify a column. |
add_fuzzy_match |
Adds a fuzzy matching node to join data on approximate string matches. |
add_google_analytics_reader |
Adds a node that reads from a Google Analytics 4 property. |
add_graph_solver |
Adds a node that solves graph-like problems within the data. |
add_group_by |
Adds a group-by aggregation node to the graph. |
add_include_cols |
Adds columns to both the input and output column lists. |
add_initial_node_analysis |
Adds a data exploration/analysis node based on a node promise. |
add_join |
Adds a join node to combine two data streams based on key columns. |
add_kafka_source |
Adds a node to read data from a Kafka or Redpanda topic. |
add_manual_input |
Adds a node for manual data entry. |
add_missing_user_defined_node |
Adds a placeholder for a custom node that cannot be loaded on this machine. |
add_node_promise |
Adds a placeholder node to the graph that is not yet fully configured. |
add_node_step |
The core method for adding or updating a node in the graph. |
add_node_to_starting_list |
Adds a node to the list of starting nodes for the flow if not already present. |
add_nodes_to_group |
Add nodes to an existing group and refit its bounds. |
add_output |
Adds an output node to write the final data to a destination. |
add_pivot |
Adds a pivot node to the graph. |
add_polars_code |
Adds a node that executes custom Polars code. |
add_python_script |
Adds a node that executes Python code on a kernel container. |
add_random_split |
Adds a node that randomly partitions rows into N labeled outputs. |
add_read |
Adds a node to read data from a local file (e.g., CSV, Parquet, Excel). |
add_record_count |
Adds a filter node to the graph. |
add_record_id |
Adds a node to create a new column with a unique ID for each record. |
add_rest_api_reader |
Adds a node that reads from a REST API. |
add_run_flow |
Adds a node that executes a catalog-registered flow as a subflow. |
add_sample |
Adds a node to take a random or top-N sample of the data. |
add_select |
Adds a node to select, rename, reorder, or drop columns. |
add_sort |
Adds a node to sort the data based on one or more columns. |
add_sql_query |
Adds a node that executes a SQL query against connected data sources. |
add_sql_source |
Adds a node that reads data from a SQL source. |
add_text_to_rows |
Adds a node that splits cell values into multiple rows. |
add_train_model |
Adds a Train Model node. |
add_union |
Adds a union node to combine multiple data streams. |
add_unique |
Adds a node to find and remove duplicate rows. |
add_unpivot |
Adds an unpivot node to the graph. |
add_user_defined_node |
Adds a user-defined custom node to the graph. |
add_wait_for |
Adds a Wait For node — passes the left input through and waits on the right. |
add_window_functions |
Adds a window-functions node (rolling, cumulative, rank, tile). |
apply_layout |
Calculates and applies a layered layout to all nodes in the graph. |
assign_node_to_named_group |
Assign a node to a group identified by name, creating it if absent (find-or-create). |
cancel |
Cancels an ongoing graph execution. |
capture_history_if_changed |
Capture history only if the flow state actually changed. |
capture_history_snapshot |
Capture the current state before a change for undo support. |
check_flow_laziness |
Check whether the flow supports lazy execution for virtual tables. |
close_flow |
Performs cleanup operations, such as clearing node caches. |
copy_node |
Creates a copy of an existing node. |
create_group |
Create a visual group. Organizational only. |
delete_group |
Remove a group box (ungroup). Members and sub-groups lift up one level. |
delete_node |
Deletes a node from the graph and updates all its connections. |
generate_code |
Generates code for the flow graph. |
get_frontend_data |
Formats the graph structure into a JSON-like dictionary for a specific legacy frontend. |
get_history_state |
Get the current state of the history system. |
get_implicit_starter_nodes |
Finds nodes that can act as starting points but are not explicitly defined as such. |
get_node |
Retrieves a node from the graph by its ID. |
get_node_data |
Retrieves all data needed to render a node in the UI. |
get_node_storage |
Serializes the entire graph's state into a storable format. |
get_nodes_overview |
Gets a list of dictionary representations for all nodes in the graph. |
get_run_info |
Gets a summary of the most recent graph execution. |
get_vue_flow_input |
Formats the graph's nodes and edges into a schema suitable for the VueFlow frontend. |
has_unsaved_changes |
Return True if the flow has changed since the last save point. |
mark_as_saved |
Mark the current flow state as the saved baseline (for dirty tracking). |
print_tree |
Print flow_graph as a visual tree structure, showing the DAG relationships with ASCII art. |
redo |
Redo the last undone action. |
release_run |
Release the single-run slot claimed by try_claim_run (idempotent). |
remove_from_output_cols |
Removes specified columns from the list of expected output columns. |
remove_nodes_from_group |
Remove nodes from whatever group they belong to; prune groups left empty. |
reset |
Forces a deep reset on all nodes in the graph. |
restore_from_snapshot |
Clear current state and rebuild from a snapshot. |
restore_groups |
Replace the runtime group registry (used by open_flow and restore_from_snapshot). |
run_graph |
Executes the entire data flow graph from start to finish. |
save_flow |
Saves the current state of the flow graph to a file. |
set_group_bounds |
Persist group box bounds (used together with set_node_positions on drag/resize). |
set_node_positions |
Persist dragged node positions (absolute canvas coordinates) onto setting_input. |
trigger_fetch_node |
Executes a specific node in the graph by its ID. |
try_claim_run |
Atomically claim the flow's single-run slot; False when a run is already in flight. |
undo |
Undo the last action by restoring to the previous state. |
update_group |
Rename / recolor / move / resize / collapse a group box. |
Attributes:
| Name | Type | Description |
|---|---|---|
execution_location |
ExecutionLocationsLiteral
|
Gets the current execution location. |
execution_mode |
ExecutionModeLiteral
|
Gets the current execution mode ('Development' or 'Performance'). |
flow_id |
int
|
Gets the unique identifier of the flow. |
graph_has_functions |
bool
|
Checks if the graph has any nodes. |
graph_has_input_data |
bool
|
Checks if the graph has an initial input data source. |
node_connections |
list[tuple[int, int]]
|
Computes and returns a list of all connections in the graph. |
nodes |
list[FlowNode]
|
Gets a list of all FlowNode objects in the graph. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 2252 2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 2327 2328 2329 2330 2331 2332 2333 2334 2335 2336 2337 2338 2339 2340 2341 2342 2343 2344 2345 2346 2347 2348 2349 2350 2351 2352 2353 2354 2355 2356 2357 2358 2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 2398 2399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 2443 2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 2467 2468 2469 2470 2471 2472 2473 2474 2475 2476 2477 2478 2479 2480 2481 2482 2483 2484 2485 2486 2487 2488 2489 2490 2491 2492 2493 2494 2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 2579 2580 2581 2582 2583 2584 2585 2586 2587 2588 2589 2590 2591 2592 2593 2594 2595 2596 2597 2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 2627 2628 2629 2630 2631 2632 2633 2634 2635 2636 2637 2638 2639 2640 2641 2642 2643 2644 2645 2646 2647 2648 2649 2650 2651 2652 2653 2654 2655 2656 2657 2658 2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 2691 2692 2693 2694 2695 2696 2697 2698 2699 2700 2701 2702 2703 2704 2705 2706 2707 2708 2709 2710 2711 2712 2713 2714 2715 2716 2717 2718 2719 2720 2721 2722 2723 2724 2725 2726 2727 2728 2729 2730 2731 2732 2733 2734 2735 2736 2737 2738 2739 2740 2741 2742 2743 2744 2745 2746 2747 2748 2749 2750 2751 2752 2753 2754 2755 2756 2757 2758 2759 2760 2761 2762 2763 2764 2765 2766 2767 2768 2769 2770 2771 2772 2773 2774 2775 2776 2777 2778 2779 2780 2781 2782 2783 2784 2785 2786 2787 2788 2789 2790 2791 2792 2793 2794 2795 2796 2797 2798 2799 2800 2801 2802 2803 2804 2805 2806 2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 2819 2820 2821 2822 2823 2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 2843 2844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 2861 2862 2863 2864 2865 2866 2867 2868 2869 2870 2871 2872 2873 2874 2875 2876 2877 2878 2879 2880 2881 2882 2883 2884 2885 2886 2887 2888 2889 2890 2891 2892 2893 2894 2895 2896 2897 2898 2899 2900 2901 2902 2903 2904 2905 2906 2907 2908 2909 2910 2911 2912 2913 2914 2915 2916 2917 2918 2919 2920 2921 2922 2923 2924 2925 2926 2927 2928 2929 2930 2931 2932 2933 2934 2935 2936 2937 2938 2939 2940 2941 2942 2943 2944 2945 2946 2947 2948 2949 2950 2951 2952 2953 2954 2955 2956 2957 2958 2959 2960 2961 2962 2963 2964 2965 2966 2967 2968 2969 2970 2971 2972 2973 2974 2975 2976 2977 2978 2979 2980 2981 2982 2983 2984 2985 2986 2987 2988 2989 2990 2991 2992 2993 2994 2995 2996 2997 2998 2999 3000 3001 3002 3003 3004 3005 3006 3007 3008 3009 3010 3011 3012 3013 3014 3015 3016 3017 3018 3019 3020 3021 3022 3023 3024 3025 3026 3027 3028 3029 3030 3031 3032 3033 3034 3035 3036 3037 3038 3039 3040 3041 3042 3043 3044 3045 3046 3047 3048 3049 3050 3051 3052 3053 3054 3055 3056 3057 3058 3059 3060 3061 3062 3063 3064 3065 3066 3067 3068 3069 3070 3071 3072 3073 3074 3075 3076 3077 3078 3079 3080 3081 3082 3083 3084 3085 3086 3087 3088 3089 3090 3091 3092 3093 3094 3095 3096 3097 3098 3099 3100 3101 3102 3103 3104 3105 3106 3107 3108 3109 3110 3111 3112 3113 3114 3115 3116 3117 3118 3119 3120 3121 3122 3123 3124 3125 3126 3127 3128 3129 3130 3131 3132 3133 3134 3135 3136 3137 3138 3139 3140 3141 3142 3143 3144 3145 3146 3147 3148 3149 3150 3151 3152 3153 3154 3155 3156 3157 3158 3159 3160 3161 3162 3163 3164 3165 3166 3167 3168 3169 3170 3171 3172 3173 3174 3175 3176 3177 3178 3179 3180 3181 3182 3183 3184 3185 3186 3187 3188 3189 3190 3191 3192 3193 3194 3195 3196 3197 3198 3199 3200 3201 3202 3203 3204 3205 3206 3207 3208 3209 3210 3211 3212 3213 3214 3215 3216 3217 3218 3219 3220 3221 3222 3223 3224 3225 3226 3227 3228 3229 3230 3231 3232 3233 3234 3235 3236 3237 3238 3239 3240 3241 3242 3243 3244 3245 3246 3247 3248 3249 3250 3251 3252 3253 3254 3255 3256 3257 3258 3259 3260 3261 3262 3263 3264 3265 3266 3267 3268 3269 3270 3271 3272 3273 3274 3275 3276 3277 3278 3279 3280 3281 3282 3283 3284 3285 3286 3287 3288 3289 3290 3291 3292 3293 3294 3295 3296 3297 3298 3299 3300 3301 3302 3303 3304 3305 3306 3307 3308 3309 3310 3311 3312 3313 3314 3315 3316 3317 3318 3319 3320 3321 3322 3323 3324 3325 3326 3327 3328 3329 3330 3331 3332 3333 3334 3335 3336 3337 3338 3339 3340 3341 3342 3343 3344 3345 3346 3347 3348 3349 3350 3351 3352 3353 3354 3355 3356 3357 3358 3359 3360 3361 3362 3363 3364 3365 3366 3367 3368 3369 3370 3371 3372 3373 3374 3375 3376 3377 3378 3379 3380 3381 3382 3383 3384 3385 3386 3387 3388 3389 3390 3391 3392 3393 3394 3395 3396 3397 3398 3399 3400 3401 3402 3403 3404 3405 3406 3407 3408 3409 3410 3411 3412 3413 3414 3415 3416 3417 3418 3419 3420 3421 3422 3423 3424 3425 3426 3427 3428 3429 3430 3431 3432 3433 3434 3435 3436 3437 3438 3439 3440 3441 3442 3443 3444 3445 3446 3447 3448 3449 3450 3451 3452 3453 3454 3455 3456 3457 3458 3459 3460 3461 3462 3463 3464 3465 3466 3467 3468 3469 3470 3471 3472 3473 3474 3475 3476 3477 3478 3479 3480 3481 3482 3483 3484 3485 3486 3487 3488 3489 3490 3491 3492 3493 3494 3495 3496 3497 3498 3499 3500 3501 3502 3503 3504 3505 3506 3507 3508 3509 3510 3511 3512 3513 3514 3515 3516 3517 3518 3519 3520 3521 3522 3523 3524 3525 3526 3527 3528 3529 3530 3531 3532 3533 3534 3535 3536 3537 3538 3539 3540 3541 3542 3543 3544 3545 3546 3547 3548 3549 3550 3551 3552 3553 3554 3555 3556 3557 3558 3559 3560 3561 3562 3563 3564 3565 3566 3567 3568 3569 3570 3571 3572 3573 3574 3575 3576 3577 3578 3579 3580 3581 3582 3583 3584 3585 3586 3587 3588 3589 3590 3591 3592 3593 3594 3595 3596 3597 3598 3599 3600 3601 3602 3603 3604 3605 3606 3607 3608 3609 3610 3611 3612 3613 3614 3615 3616 3617 3618 3619 3620 3621 3622 3623 3624 3625 3626 3627 3628 3629 3630 3631 3632 3633 3634 3635 3636 3637 3638 3639 3640 3641 3642 3643 3644 3645 3646 3647 3648 3649 3650 3651 3652 3653 3654 3655 3656 3657 3658 3659 3660 3661 3662 3663 3664 3665 3666 3667 3668 3669 3670 3671 3672 3673 3674 3675 3676 3677 3678 3679 3680 3681 3682 3683 3684 3685 3686 3687 3688 3689 3690 3691 3692 3693 3694 3695 3696 3697 3698 3699 3700 3701 3702 3703 3704 3705 3706 3707 3708 3709 3710 3711 3712 3713 3714 3715 3716 3717 3718 3719 3720 3721 3722 3723 3724 3725 3726 3727 3728 3729 3730 3731 3732 3733 3734 3735 3736 3737 3738 3739 3740 3741 3742 3743 3744 3745 3746 3747 3748 3749 3750 3751 3752 3753 3754 3755 3756 3757 3758 3759 3760 3761 3762 3763 3764 3765 3766 3767 3768 3769 3770 3771 3772 3773 3774 3775 3776 3777 3778 3779 3780 3781 3782 3783 3784 3785 3786 3787 3788 3789 3790 3791 3792 3793 3794 3795 3796 3797 3798 3799 3800 3801 3802 3803 3804 3805 3806 3807 3808 3809 3810 3811 3812 3813 3814 3815 3816 3817 3818 3819 3820 3821 3822 3823 3824 3825 3826 3827 3828 3829 3830 3831 3832 3833 3834 3835 3836 3837 3838 3839 3840 3841 3842 3843 3844 3845 3846 3847 3848 3849 3850 3851 3852 3853 3854 3855 3856 3857 3858 3859 3860 3861 3862 3863 3864 3865 3866 3867 3868 3869 3870 3871 3872 3873 3874 3875 3876 3877 3878 3879 3880 3881 3882 3883 3884 3885 3886 3887 3888 3889 3890 3891 3892 3893 3894 3895 3896 3897 3898 3899 3900 3901 3902 3903 3904 3905 3906 3907 3908 3909 3910 3911 3912 3913 3914 3915 3916 3917 3918 3919 3920 3921 3922 3923 3924 3925 3926 3927 3928 3929 3930 3931 3932 3933 3934 3935 3936 3937 3938 3939 3940 3941 3942 3943 3944 3945 3946 3947 3948 3949 3950 3951 3952 3953 3954 3955 3956 3957 3958 3959 3960 3961 3962 3963 3964 3965 3966 3967 3968 3969 3970 3971 3972 3973 3974 3975 3976 3977 3978 3979 3980 3981 3982 3983 3984 3985 3986 3987 3988 3989 3990 3991 3992 3993 3994 3995 3996 3997 3998 3999 4000 4001 4002 4003 4004 4005 4006 4007 4008 4009 4010 4011 4012 4013 4014 4015 4016 4017 4018 4019 4020 4021 4022 4023 4024 4025 4026 4027 4028 4029 4030 4031 4032 4033 4034 4035 4036 4037 4038 4039 4040 4041 4042 4043 4044 4045 4046 4047 4048 4049 4050 4051 4052 4053 4054 4055 4056 4057 4058 4059 4060 4061 4062 4063 4064 4065 4066 4067 4068 4069 4070 4071 4072 4073 4074 4075 4076 4077 4078 4079 4080 4081 4082 4083 4084 4085 4086 4087 4088 4089 4090 4091 4092 4093 4094 4095 4096 4097 4098 4099 4100 4101 4102 4103 4104 4105 4106 4107 4108 4109 4110 4111 4112 4113 4114 4115 4116 4117 4118 4119 4120 4121 4122 4123 4124 4125 4126 4127 4128 4129 4130 4131 4132 4133 4134 4135 4136 4137 4138 4139 4140 4141 4142 4143 4144 4145 4146 4147 4148 4149 4150 4151 4152 4153 4154 4155 4156 4157 4158 4159 4160 4161 4162 4163 4164 4165 4166 4167 4168 4169 4170 4171 4172 4173 4174 4175 4176 4177 4178 4179 4180 4181 4182 4183 4184 4185 4186 4187 4188 4189 4190 4191 4192 4193 4194 4195 4196 4197 4198 4199 4200 4201 4202 4203 4204 4205 4206 4207 4208 4209 4210 4211 4212 4213 4214 4215 4216 4217 4218 4219 4220 4221 4222 4223 4224 4225 4226 4227 4228 4229 4230 4231 4232 4233 4234 4235 4236 4237 4238 4239 4240 4241 4242 4243 4244 4245 4246 4247 4248 4249 4250 4251 4252 4253 4254 4255 4256 4257 4258 4259 4260 4261 4262 4263 4264 4265 4266 4267 4268 4269 4270 4271 4272 4273 4274 4275 4276 4277 4278 4279 4280 4281 4282 4283 4284 4285 4286 4287 4288 4289 4290 4291 4292 4293 4294 4295 4296 4297 4298 4299 4300 4301 4302 4303 4304 4305 4306 4307 4308 4309 4310 4311 4312 4313 4314 4315 4316 4317 4318 4319 4320 4321 4322 4323 4324 4325 4326 4327 4328 4329 4330 4331 4332 4333 4334 4335 4336 4337 4338 4339 4340 4341 4342 4343 4344 4345 4346 4347 4348 4349 4350 4351 4352 4353 4354 4355 4356 4357 4358 4359 4360 4361 4362 4363 4364 4365 4366 4367 4368 4369 4370 4371 4372 4373 4374 4375 4376 4377 4378 4379 4380 4381 4382 4383 4384 4385 4386 4387 4388 4389 4390 4391 4392 4393 4394 4395 4396 4397 4398 4399 4400 4401 4402 4403 4404 4405 4406 4407 4408 4409 4410 4411 4412 4413 4414 4415 4416 4417 4418 4419 4420 4421 4422 4423 4424 4425 4426 4427 4428 4429 4430 4431 4432 4433 4434 4435 4436 4437 4438 4439 4440 4441 4442 4443 4444 4445 4446 4447 4448 4449 4450 4451 4452 4453 4454 4455 4456 4457 4458 4459 4460 4461 4462 4463 4464 4465 4466 4467 4468 4469 4470 4471 4472 4473 4474 4475 4476 4477 4478 4479 4480 4481 4482 4483 4484 4485 4486 4487 4488 4489 4490 4491 4492 4493 4494 4495 4496 4497 4498 4499 4500 4501 4502 4503 4504 4505 4506 4507 4508 4509 4510 4511 4512 4513 4514 4515 4516 4517 4518 4519 4520 4521 4522 4523 4524 4525 4526 4527 4528 4529 4530 4531 4532 4533 4534 4535 4536 4537 4538 4539 4540 4541 4542 4543 4544 4545 4546 4547 4548 4549 4550 4551 4552 4553 4554 4555 4556 4557 4558 4559 4560 4561 4562 4563 4564 4565 4566 4567 4568 4569 4570 4571 4572 4573 4574 4575 4576 4577 4578 4579 4580 4581 4582 4583 4584 4585 4586 4587 4588 4589 4590 4591 4592 4593 4594 4595 4596 4597 4598 4599 4600 4601 4602 4603 4604 4605 4606 4607 4608 4609 4610 4611 4612 4613 4614 4615 4616 4617 4618 4619 4620 4621 4622 4623 4624 4625 4626 4627 4628 4629 4630 4631 4632 4633 4634 4635 4636 4637 4638 4639 4640 4641 4642 4643 4644 4645 4646 4647 4648 4649 4650 4651 4652 4653 4654 4655 4656 4657 4658 4659 4660 4661 4662 4663 4664 4665 4666 4667 4668 4669 4670 4671 4672 4673 4674 4675 4676 4677 4678 4679 4680 4681 4682 4683 4684 4685 4686 4687 4688 4689 4690 4691 4692 4693 4694 4695 4696 4697 4698 4699 4700 4701 4702 4703 4704 4705 4706 4707 4708 4709 4710 4711 4712 4713 4714 4715 4716 4717 4718 4719 4720 4721 4722 4723 4724 4725 4726 4727 4728 4729 4730 4731 4732 4733 4734 4735 4736 4737 4738 4739 4740 4741 4742 4743 4744 4745 4746 4747 4748 4749 4750 4751 4752 4753 4754 4755 4756 4757 4758 4759 4760 4761 4762 4763 4764 4765 4766 4767 4768 4769 4770 4771 4772 4773 4774 4775 4776 4777 4778 4779 4780 4781 4782 4783 4784 4785 4786 4787 4788 4789 4790 4791 4792 4793 4794 4795 4796 4797 4798 4799 4800 4801 4802 4803 4804 4805 4806 4807 4808 4809 4810 4811 4812 4813 4814 4815 4816 4817 4818 4819 4820 4821 4822 4823 4824 4825 4826 4827 4828 4829 4830 4831 4832 4833 4834 4835 4836 4837 4838 4839 4840 4841 4842 4843 4844 4845 4846 4847 4848 4849 4850 4851 4852 4853 4854 4855 4856 4857 4858 4859 4860 4861 4862 4863 4864 4865 4866 4867 4868 4869 4870 4871 4872 4873 4874 4875 4876 4877 4878 4879 4880 4881 4882 4883 4884 4885 4886 4887 4888 4889 4890 4891 4892 4893 4894 4895 4896 4897 4898 4899 4900 4901 4902 4903 4904 4905 4906 4907 4908 4909 4910 4911 4912 4913 4914 4915 4916 4917 4918 4919 4920 4921 4922 4923 4924 4925 4926 4927 4928 4929 4930 4931 4932 4933 4934 4935 4936 4937 4938 4939 4940 4941 4942 4943 4944 4945 4946 4947 4948 4949 4950 4951 4952 4953 4954 4955 4956 4957 4958 4959 4960 4961 4962 4963 4964 4965 4966 4967 4968 4969 4970 4971 4972 4973 4974 4975 4976 4977 4978 4979 4980 4981 4982 4983 4984 4985 4986 4987 4988 4989 4990 4991 4992 4993 4994 4995 4996 4997 4998 4999 5000 5001 5002 5003 5004 5005 5006 5007 5008 5009 5010 5011 5012 5013 5014 5015 5016 5017 5018 5019 5020 5021 5022 5023 5024 5025 5026 5027 5028 5029 5030 5031 5032 5033 5034 5035 5036 5037 5038 5039 5040 5041 5042 5043 5044 5045 5046 5047 5048 5049 5050 5051 5052 5053 5054 5055 5056 5057 5058 5059 5060 5061 5062 5063 5064 5065 5066 5067 5068 5069 5070 5071 5072 5073 5074 5075 5076 5077 5078 5079 5080 5081 5082 5083 5084 5085 5086 5087 5088 5089 5090 5091 5092 5093 5094 5095 5096 5097 5098 5099 5100 5101 5102 5103 5104 5105 5106 5107 5108 5109 5110 5111 5112 5113 5114 5115 5116 5117 5118 5119 5120 5121 5122 5123 5124 5125 5126 5127 5128 5129 5130 5131 5132 5133 5134 5135 5136 5137 5138 5139 5140 5141 5142 5143 5144 5145 5146 5147 5148 5149 5150 5151 5152 5153 5154 5155 5156 5157 5158 5159 5160 5161 5162 5163 5164 5165 5166 5167 5168 5169 5170 5171 5172 5173 5174 5175 5176 5177 5178 5179 5180 5181 5182 5183 5184 5185 5186 5187 5188 5189 5190 5191 5192 5193 5194 5195 5196 5197 5198 5199 5200 5201 5202 5203 5204 5205 5206 5207 5208 5209 5210 5211 5212 5213 5214 5215 5216 5217 5218 5219 5220 5221 5222 5223 5224 5225 5226 5227 5228 5229 5230 5231 5232 5233 5234 5235 5236 5237 5238 5239 5240 5241 5242 5243 5244 5245 5246 5247 5248 5249 5250 5251 5252 5253 5254 5255 5256 5257 5258 5259 5260 5261 5262 5263 5264 5265 5266 5267 5268 5269 5270 5271 5272 5273 5274 5275 5276 5277 5278 5279 5280 5281 5282 5283 5284 5285 5286 5287 5288 5289 5290 5291 5292 5293 5294 5295 5296 5297 5298 5299 5300 5301 5302 5303 5304 5305 5306 5307 5308 5309 5310 5311 5312 5313 5314 5315 5316 5317 5318 5319 5320 5321 5322 5323 5324 5325 5326 5327 5328 5329 5330 5331 5332 5333 5334 5335 5336 5337 5338 5339 5340 5341 5342 5343 5344 5345 5346 5347 5348 5349 5350 5351 5352 5353 5354 5355 5356 5357 5358 5359 5360 5361 5362 5363 5364 5365 5366 5367 5368 5369 5370 5371 5372 5373 5374 5375 5376 5377 5378 5379 5380 5381 5382 5383 5384 5385 5386 5387 5388 5389 5390 5391 5392 5393 5394 5395 5396 5397 5398 5399 5400 5401 5402 5403 5404 5405 5406 5407 5408 5409 5410 5411 5412 5413 5414 5415 5416 5417 5418 5419 5420 5421 5422 5423 5424 5425 5426 5427 5428 5429 5430 5431 5432 5433 5434 5435 5436 5437 5438 5439 5440 5441 5442 5443 5444 5445 5446 5447 5448 5449 5450 5451 5452 5453 5454 5455 5456 5457 5458 5459 5460 5461 5462 5463 5464 5465 5466 5467 5468 5469 5470 5471 5472 5473 5474 5475 5476 5477 5478 5479 5480 5481 5482 5483 5484 5485 5486 5487 5488 5489 5490 5491 5492 5493 5494 5495 5496 5497 5498 5499 5500 5501 5502 5503 5504 5505 5506 5507 5508 5509 5510 5511 5512 5513 5514 5515 5516 5517 5518 5519 5520 5521 5522 5523 5524 5525 5526 5527 5528 5529 5530 5531 5532 5533 5534 5535 5536 5537 5538 5539 5540 5541 5542 5543 5544 5545 5546 5547 5548 5549 5550 5551 5552 5553 5554 5555 5556 5557 5558 5559 5560 5561 5562 5563 5564 5565 5566 5567 5568 5569 5570 5571 5572 5573 5574 5575 5576 5577 5578 5579 5580 5581 5582 5583 5584 5585 5586 5587 5588 5589 5590 5591 5592 5593 5594 5595 5596 5597 5598 5599 5600 5601 5602 5603 5604 5605 5606 5607 5608 5609 5610 5611 5612 5613 5614 5615 5616 5617 5618 5619 5620 5621 5622 5623 5624 5625 5626 5627 5628 5629 5630 5631 5632 5633 5634 5635 5636 5637 5638 5639 5640 5641 5642 5643 5644 5645 5646 5647 5648 5649 5650 5651 5652 5653 5654 5655 5656 5657 5658 5659 5660 5661 5662 5663 5664 5665 5666 5667 5668 5669 5670 5671 5672 5673 5674 5675 5676 5677 5678 5679 5680 5681 5682 5683 5684 5685 5686 5687 5688 5689 5690 5691 5692 5693 5694 5695 5696 5697 5698 5699 5700 5701 5702 5703 5704 5705 5706 5707 5708 5709 5710 5711 5712 5713 5714 5715 5716 5717 5718 5719 5720 5721 5722 5723 5724 5725 5726 5727 5728 5729 5730 5731 5732 5733 5734 5735 5736 5737 5738 5739 5740 5741 5742 5743 5744 5745 5746 5747 5748 5749 5750 5751 5752 5753 5754 5755 5756 5757 5758 5759 5760 5761 5762 5763 5764 5765 5766 5767 5768 5769 5770 5771 5772 5773 5774 5775 5776 5777 5778 5779 5780 5781 5782 5783 5784 5785 5786 5787 5788 5789 5790 5791 5792 5793 5794 5795 5796 5797 5798 5799 5800 5801 5802 5803 5804 5805 5806 5807 5808 5809 5810 5811 5812 5813 5814 5815 5816 5817 5818 5819 5820 5821 5822 5823 5824 5825 5826 5827 5828 5829 5830 5831 5832 5833 5834 5835 5836 5837 5838 5839 5840 5841 5842 5843 5844 5845 5846 5847 5848 5849 5850 5851 5852 5853 5854 5855 5856 5857 5858 5859 5860 5861 5862 5863 5864 5865 5866 5867 5868 5869 5870 5871 5872 5873 5874 5875 5876 5877 5878 5879 5880 5881 5882 5883 5884 5885 5886 5887 5888 5889 5890 5891 5892 5893 5894 5895 5896 5897 5898 5899 5900 5901 5902 5903 5904 5905 5906 5907 5908 5909 5910 5911 5912 5913 5914 5915 5916 5917 5918 5919 5920 5921 5922 5923 5924 5925 5926 5927 5928 5929 5930 5931 5932 5933 5934 5935 5936 5937 5938 5939 5940 5941 5942 5943 5944 5945 5946 5947 5948 5949 5950 5951 5952 5953 5954 5955 5956 5957 5958 5959 5960 5961 5962 5963 5964 5965 5966 5967 5968 5969 5970 5971 5972 5973 5974 5975 5976 5977 5978 5979 5980 5981 5982 5983 5984 5985 5986 5987 5988 5989 5990 5991 5992 5993 5994 5995 5996 5997 5998 5999 6000 6001 6002 6003 6004 6005 6006 6007 6008 6009 6010 6011 6012 6013 6014 6015 6016 6017 6018 6019 6020 6021 6022 6023 6024 6025 6026 6027 6028 6029 6030 6031 6032 6033 6034 6035 6036 6037 6038 6039 6040 6041 6042 6043 6044 6045 6046 6047 6048 6049 6050 6051 6052 6053 6054 6055 6056 6057 6058 6059 6060 6061 6062 6063 6064 6065 6066 6067 6068 6069 6070 6071 6072 6073 6074 6075 6076 6077 6078 6079 6080 6081 6082 6083 6084 6085 6086 6087 6088 6089 6090 6091 6092 6093 6094 6095 6096 6097 6098 6099 6100 6101 6102 6103 6104 6105 6106 6107 6108 6109 6110 6111 6112 6113 6114 6115 6116 6117 6118 6119 6120 6121 6122 6123 6124 6125 6126 6127 6128 6129 6130 6131 6132 6133 6134 6135 6136 6137 6138 6139 6140 6141 6142 6143 6144 6145 6146 6147 6148 6149 6150 6151 6152 6153 6154 6155 6156 6157 6158 6159 6160 6161 6162 6163 6164 6165 6166 6167 6168 6169 6170 6171 6172 6173 6174 6175 6176 6177 6178 6179 6180 6181 6182 6183 6184 6185 6186 6187 6188 6189 6190 6191 6192 6193 6194 6195 6196 6197 6198 6199 6200 6201 6202 6203 6204 6205 6206 6207 6208 6209 6210 6211 6212 6213 6214 6215 6216 6217 6218 6219 6220 6221 6222 6223 6224 6225 6226 6227 6228 6229 6230 6231 6232 6233 6234 6235 6236 6237 6238 6239 6240 6241 6242 6243 6244 6245 6246 6247 6248 6249 6250 6251 6252 6253 6254 6255 6256 6257 6258 6259 6260 6261 6262 6263 6264 6265 6266 6267 6268 6269 6270 6271 6272 6273 6274 6275 6276 6277 6278 6279 6280 6281 6282 6283 6284 6285 6286 6287 6288 6289 6290 6291 6292 6293 6294 6295 6296 6297 6298 6299 6300 6301 6302 6303 6304 6305 6306 6307 6308 6309 6310 6311 6312 6313 6314 6315 6316 6317 6318 6319 6320 6321 6322 6323 6324 6325 6326 6327 6328 6329 6330 6331 6332 6333 6334 6335 6336 6337 6338 6339 6340 6341 6342 6343 6344 6345 6346 6347 6348 6349 6350 6351 6352 6353 6354 6355 6356 6357 6358 6359 6360 6361 6362 6363 6364 6365 6366 6367 6368 6369 6370 6371 6372 6373 6374 6375 6376 6377 6378 6379 6380 6381 6382 6383 6384 6385 6386 6387 6388 6389 6390 6391 6392 6393 6394 6395 6396 6397 6398 6399 6400 6401 6402 6403 6404 6405 6406 6407 6408 6409 6410 6411 6412 6413 6414 6415 6416 6417 6418 6419 6420 6421 6422 6423 6424 6425 6426 6427 6428 6429 6430 6431 6432 6433 6434 6435 6436 6437 6438 6439 6440 6441 6442 6443 6444 6445 6446 6447 6448 6449 6450 6451 6452 6453 6454 6455 6456 6457 6458 6459 6460 6461 6462 6463 6464 6465 6466 6467 6468 6469 6470 6471 6472 6473 6474 6475 6476 6477 6478 6479 6480 6481 6482 6483 6484 6485 6486 6487 6488 6489 6490 6491 6492 6493 6494 6495 6496 6497 6498 6499 6500 6501 6502 6503 6504 6505 6506 6507 6508 6509 6510 6511 6512 6513 6514 6515 6516 6517 6518 6519 6520 6521 6522 6523 6524 6525 6526 6527 6528 6529 6530 6531 6532 6533 6534 6535 6536 6537 6538 6539 6540 6541 6542 6543 6544 6545 6546 6547 6548 6549 6550 6551 6552 6553 6554 6555 6556 6557 6558 6559 6560 6561 6562 6563 6564 6565 6566 6567 6568 6569 6570 6571 6572 6573 6574 6575 6576 6577 6578 6579 6580 6581 6582 6583 6584 6585 6586 6587 6588 6589 6590 6591 6592 6593 6594 6595 6596 6597 6598 6599 6600 6601 6602 6603 6604 6605 6606 6607 6608 6609 6610 6611 6612 6613 6614 6615 6616 6617 6618 6619 6620 6621 6622 6623 6624 6625 6626 6627 6628 6629 6630 6631 6632 6633 6634 6635 6636 6637 6638 6639 6640 6641 6642 6643 6644 6645 6646 6647 6648 6649 6650 6651 6652 6653 6654 6655 6656 6657 6658 6659 6660 6661 6662 6663 6664 6665 6666 6667 6668 6669 6670 6671 6672 6673 6674 6675 6676 6677 6678 6679 6680 6681 6682 6683 6684 6685 6686 6687 6688 6689 6690 6691 6692 6693 6694 6695 6696 6697 6698 6699 6700 6701 6702 6703 6704 6705 6706 6707 6708 6709 6710 6711 6712 6713 6714 6715 6716 6717 6718 6719 6720 | |
execution_location
property
writable
Gets the current execution location.
execution_mode
property
writable
Gets the current execution mode ('Development' or 'Performance').
flow_id
property
writable
Gets the unique identifier of the flow.
graph_has_functions
property
Checks if the graph has any nodes.
graph_has_input_data
property
Checks if the graph has an initial input data source.
node_connections
property
Computes and returns a list of all connections in the graph.
Returns:
| Type | Description |
|---|---|
list[tuple[int, int]]
|
A list of tuples, where each tuple is a (source_id, target_id) pair. |
nodes
property
Gets a list of all FlowNode objects in the graph.
__init__(flow_settings, name=None, input_cols=None, output_cols=None, path_ref=None, input_flow=None, cache_results=False)
Initializes a new FlowGraph instance.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_settings
|
FlowSettings | FlowGraphConfig
|
The configuration settings for the flow. |
required |
name
|
str
|
The name of the flow. |
None
|
input_cols
|
list[str]
|
A list of input column names. |
None
|
output_cols
|
list[str]
|
A list of output column names. |
None
|
path_ref
|
str
|
An optional path to an initial data source. |
None
|
input_flow
|
Union[ParquetFile, FlowDataEngine, FlowGraph]
|
An optional existing data object to start the flow with. |
None
|
cache_results
|
bool
|
A global flag to enable or disable result caching. |
False
|
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 | |
__repr__()
Provides the official string representation of the FlowGraph instance.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2259 2260 2261 2262 | |
add_api_response(api_response)
Adds an API-response sink node.
The node is a pass-through marker: its result equals its input. When the flow is published as an HTTP API endpoint, the endpoint reads this node's result and serializes it as the response body. Behaves like an output node so its result is always materialized locally.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
api_response
|
NodeApiResponse
|
The settings for the API-response node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4372 4373 4374 4375 4376 4377 4378 4379 4380 4381 4382 4383 4384 4385 4386 4387 4388 4389 4390 4391 4392 4393 4394 4395 4396 4397 4398 4399 4400 4401 | |
add_apply_model(apply_settings)
Adds an Apply Model node.
Fetches the artifact from the catalog and asks the worker to score the
input data, returning a LazyFrame with one extra Float64 column.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
apply_settings
|
NodeApplyModel
|
Settings (model_name, optional version, output_column). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
The |
FlowGraph
|
class: |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3601 3602 3603 3604 3605 3606 3607 3608 3609 3610 3611 3612 3613 3614 3615 3616 3617 3618 3619 3620 3621 3622 3623 3624 3625 3626 3627 3628 3629 3630 3631 3632 3633 3634 3635 3636 3637 3638 3639 3640 3641 3642 3643 3644 3645 3646 3647 3648 3649 3650 3651 3652 3653 3654 3655 3656 3657 3658 3659 3660 3661 3662 3663 3664 3665 3666 3667 3668 3669 3670 3671 3672 3673 3674 3675 3676 3677 3678 3679 3680 3681 3682 3683 3684 3685 3686 3687 3688 3689 3690 3691 3692 3693 3694 3695 3696 3697 3698 3699 3700 3701 3702 3703 3704 3705 3706 3707 3708 3709 3710 3711 3712 3713 3714 3715 3716 3717 3718 3719 3720 3721 3722 | |
add_catalog_reader(node_catalog_reader)
Adds a node that reads a table from the catalog.
Resolves the catalog table by ID (or name + namespace) and reads
the materialized Parquet file. When sql_query is set, executes
the SQL against all catalog Delta tables instead.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4496 4497 4498 4499 4500 4501 4502 4503 4504 4505 4506 4507 4508 4509 | |
add_catalog_writer(node_catalog_writer)
Adds a node that writes its input to the catalog as a Delta table or virtual table.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4671 4672 4673 4674 4675 4676 4677 4678 4679 4680 4681 4682 4683 4684 4685 4686 4687 4688 4689 4690 4691 4692 4693 4694 4695 4696 | |
add_cloud_storage_reader(node_cloud_storage_reader)
Adds a cloud storage read node to the flow graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_cloud_storage_reader
|
NodeCloudStorageReader
|
The settings for the cloud storage read node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5361 5362 5363 5364 5365 5366 5367 5368 5369 5370 5371 5372 5373 5374 5375 5376 5377 5378 5379 5380 5381 5382 5383 5384 5385 5386 5387 5388 5389 5390 5391 5392 5393 | |
add_cloud_storage_writer(node_cloud_storage_writer)
Adds a node to write data to a cloud storage provider.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_cloud_storage_writer
|
NodeCloudStorageWriter
|
The settings for the cloud storage writer node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5303 5304 5305 5306 5307 5308 5309 5310 5311 5312 5313 5314 5315 5316 5317 5318 5319 5320 5321 5322 5323 5324 5325 5326 5327 5328 5329 5330 5331 5332 5333 5334 5335 5336 5337 5338 5339 5340 5341 5342 5343 5344 5345 5346 5347 5348 5349 5350 5351 5352 5353 5354 5355 5356 5357 5358 5359 | |
add_cross_join(cross_join_settings)
Adds a cross join node to the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cross_join_settings
|
NodeCrossJoin
|
The settings for the cross join operation. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3292 3293 3294 3295 3296 3297 3298 3299 3300 3301 3302 3303 3304 3305 3306 3307 3308 3309 3310 3311 3312 3313 3314 3315 3316 3317 3318 3319 3320 3321 3322 3323 3324 3325 3326 3327 3328 3329 3330 3331 3332 3333 | |
add_database_reader(node_database_reader)
Adds a node to read data from a database.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_database_reader
|
NodeDatabaseReader
|
The settings for the database reader node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4771 4772 4773 4774 4775 4776 4777 4778 4779 4780 4781 4782 4783 4784 4785 4786 4787 4788 4789 4790 4791 4792 4793 4794 4795 4796 4797 4798 4799 4800 4801 4802 4803 4804 4805 4806 4807 4808 4809 4810 4811 4812 4813 4814 4815 4816 4817 4818 4819 4820 4821 4822 4823 4824 4825 4826 4827 4828 4829 4830 4831 4832 4833 4834 4835 4836 4837 4838 4839 4840 4841 4842 4843 4844 4845 4846 4847 4848 4849 4850 4851 4852 4853 4854 4855 4856 4857 4858 4859 4860 4861 4862 4863 4864 4865 4866 4867 4868 4869 4870 4871 4872 4873 4874 4875 4876 4877 4878 4879 4880 4881 4882 4883 4884 4885 4886 4887 4888 4889 4890 4891 4892 4893 4894 4895 4896 4897 4898 4899 4900 4901 | |
add_database_writer(node_database_writer)
Adds a node to write data to a database.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_database_writer
|
NodeDatabaseWriter
|
The settings for the database writer node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4698 4699 4700 4701 4702 4703 4704 4705 4706 4707 4708 4709 4710 4711 4712 4713 4714 4715 4716 4717 4718 4719 4720 4721 4722 4723 4724 4725 4726 4727 4728 4729 4730 4731 4732 4733 4734 4735 4736 4737 4738 4739 4740 4741 4742 4743 4744 4745 4746 4747 4748 4749 4750 4751 4752 4753 4754 4755 4756 4757 4758 4759 4760 4761 4762 4763 4764 4765 4766 4767 4768 4769 | |
add_datasource(input_file)
Adds a data source node to the graph.
This method serves as a factory for creating starting nodes, handling both file-based sources and direct manual data entry.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_file
|
NodeDatasource | NodeManualInput
|
The configuration object for the data source. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5558 5559 5560 5561 5562 5563 5564 5565 5566 5567 5568 5569 5570 5571 5572 5573 5574 5575 5576 5577 5578 5579 5580 5581 5582 5583 5584 5585 5586 5587 5588 5589 5590 5591 5592 5593 5594 5595 5596 5597 5598 | |
add_dependency_on_polars_lazy_frame(lazy_frame, node_id)
Adds a special node that directly injects a Polars LazyFrame into the graph.
Note: This is intended for backend use and will not work in the UI editor.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
lazy_frame
|
LazyFrame
|
The Polars LazyFrame to inject. |
required |
node_id
|
int
|
The ID for the new node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3185 3186 3187 3188 3189 3190 3191 3192 3193 3194 3195 3196 3197 3198 3199 3200 3201 3202 3203 | |
add_dynamic_rename(settings)
Adds a node that renames many columns at once via a single rule.
Supports prefix, suffix, formula-based, and first-row renaming across all
columns, a specific list of columns, or every column of a given data type.
In first_row mode the first row is dropped from the output after its
values are promoted to column headers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
NodeDynamicRename
|
The dynamic rename configuration. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4059 4060 4061 4062 4063 4064 4065 4066 4067 4068 4069 4070 4071 4072 4073 4074 4075 4076 4077 4078 4079 4080 4081 4082 4083 4084 4085 | |
add_evaluate_model(evaluate_settings)
Adds an Evaluate Model node.
Compares the actual and predicted columns already present on the
input dataframe and emits a long-form (metric, value) frame.
Pure polars — no worker offload, no model file read.
task_type="auto" resolves the metric set from the configured
upstream Train Model node's trainer; otherwise uses the explicit
regression / classification choice from settings.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3724 3725 3726 3727 3728 3729 3730 3731 3732 3733 3734 3735 3736 3737 3738 3739 3740 3741 3742 3743 3744 3745 3746 3747 3748 3749 3750 3751 3752 3753 3754 3755 3756 3757 3758 3759 3760 3761 3762 3763 3764 3765 3766 3767 3768 3769 3770 3771 3772 3773 3774 3775 3776 3777 3778 3779 3780 3781 3782 3783 3784 3785 3786 3787 3788 3789 3790 3791 | |
add_explore_data(node_analysis)
Adds a specialized node for data exploration and visualization.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_analysis
|
NodeExploreData
|
The settings for the data exploration node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2943 2944 2945 2946 2947 2948 2949 2950 2951 2952 2953 2954 2955 2956 2957 2958 2959 2960 2961 2962 2963 2964 2965 2966 2967 2968 2969 2970 2971 2972 2973 | |
add_external_source(external_source_input)
Adds a node for a custom external data source.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
external_source_input
|
NodeExternalSource
|
The settings for the external source node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5395 5396 5397 5398 5399 5400 5401 5402 5403 5404 5405 5406 5407 5408 5409 5410 5411 5412 5413 5414 5415 5416 5417 5418 5419 5420 5421 5422 5423 5424 5425 5426 5427 5428 5429 5430 5431 5432 5433 5434 5435 5436 5437 5438 5439 5440 5441 5442 5443 5444 5445 5446 5447 5448 5449 5450 5451 5452 5453 5454 5455 5456 5457 5458 5459 5460 5461 5462 5463 5464 5465 5466 5467 | |
add_filter(filter_settings)
Adds a filter node to the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
filter_settings
|
NodeFilter
|
The settings for the filter operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3008 3009 3010 3011 3012 3013 3014 3015 3016 3017 3018 3019 3020 3021 3022 3023 3024 3025 3026 3027 3028 3029 3030 3031 3032 3033 3034 3035 3036 3037 3038 3039 3040 3041 3042 3043 3044 3045 3046 | |
add_flow_input(settings)
Adds a named subflow-input placeholder source.
Standalone runs serve the optional sample data (empty frame otherwise);
a parent run_flow node overwrites node.function with real data.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5610 5611 5612 5613 5614 5615 5616 5617 5618 5619 5620 5621 5622 5623 5624 5625 5626 5627 5628 5629 5630 5631 5632 5633 5634 5635 5636 5637 5638 5639 5640 5641 5642 5643 5644 5645 5646 5647 5648 | |
add_flow_output(settings)
Adds a named subflow-output sink (passthrough, always materialized).
When this flow runs inside another flow via a run_flow node, the parent reads this node's result as one of the subflow's outputs.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4403 4404 4405 4406 4407 4408 4409 4410 4411 4412 4413 4414 4415 4416 4417 4418 4419 4420 4421 4422 4423 4424 4425 4426 4427 4428 4429 4430 4431 4432 4433 4434 4435 4436 4437 | |
add_formula(function_settings)
Adds a node that applies a formula to create or modify a column.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
function_settings
|
NodeFormula
|
The settings for the formula operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3249 3250 3251 3252 3253 3254 3255 3256 3257 3258 3259 3260 3261 3262 3263 3264 3265 3266 3267 3268 3269 3270 3271 3272 3273 3274 3275 3276 3277 3278 3279 3280 3281 3282 3283 3284 3285 3286 3287 3288 3289 3290 | |
add_fuzzy_match(fuzzy_settings)
Adds a fuzzy matching node to join data on approximate string matches.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fuzzy_settings
|
NodeFuzzyMatch
|
The settings for the fuzzy match operation. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3825 3826 3827 3828 3829 3830 3831 3832 3833 3834 3835 3836 3837 3838 3839 3840 3841 3842 3843 3844 3845 3846 3847 3848 3849 3850 3851 3852 3853 3854 3855 3856 3857 3858 3859 3860 3861 3862 3863 3864 3865 3866 3867 3868 3869 3870 3871 3872 3873 3874 3875 3876 3877 | |
add_google_analytics_reader(node_ga_reader)
Adds a node that reads from a Google Analytics 4 property.
The actual API fetch (OAuth token refresh, run_report calls,
pagination) is offloaded to the worker via ExternalGoogleAnalyticsFetcher,
so the core's event loop stays responsive. The schema_callback is
derived locally from the selected metrics/dimensions — no network call
is made during schema prediction, keeping downstream nodes lazy.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5077 5078 5079 5080 5081 5082 5083 5084 5085 5086 5087 5088 5089 5090 5091 5092 5093 5094 5095 5096 5097 5098 5099 5100 5101 5102 5103 5104 5105 5106 5107 5108 5109 5110 5111 5112 5113 5114 5115 5116 5117 5118 5119 5120 5121 5122 5123 5124 5125 5126 5127 5128 5129 5130 5131 5132 5133 5134 5135 5136 5137 5138 5139 5140 5141 5142 5143 5144 5145 5146 5147 5148 5149 5150 5151 5152 5153 5154 5155 5156 5157 5158 5159 5160 5161 5162 5163 5164 5165 5166 5167 5168 5169 5170 5171 5172 5173 5174 5175 5176 5177 5178 5179 5180 5181 5182 5183 5184 5185 5186 5187 5188 5189 5190 5191 5192 5193 5194 5195 5196 5197 5198 5199 5200 5201 5202 5203 5204 5205 5206 5207 5208 5209 5210 5211 5212 5213 5214 5215 5216 5217 5218 5219 5220 | |
add_graph_solver(graph_solver_settings)
Adds a node that solves graph-like problems within the data.
This node can be used for operations like finding network paths, calculating connected components, or performing other graph algorithms on relational data that represents nodes and edges.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph_solver_settings
|
NodeGraphSolver
|
The settings object defining the graph inputs and the specific algorithm to apply. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3225 3226 3227 3228 3229 3230 3231 3232 3233 3234 3235 3236 3237 3238 3239 3240 3241 3242 3243 3244 3245 3246 3247 | |
add_group_by(group_by_settings)
Adds a group-by aggregation node to the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
group_by_settings
|
NodeGroupBy
|
The settings for the group-by operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2975 2976 2977 2978 2979 2980 2981 2982 2983 2984 2985 2986 2987 2988 2989 2990 2991 2992 2993 2994 2995 2996 2997 2998 2999 3000 3001 3002 3003 3004 3005 3006 | |
add_include_cols(include_columns)
Adds columns to both the input and output column lists.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
include_columns
|
list[str]
|
A list of column names to include. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4308 4309 4310 4311 4312 4313 4314 4315 4316 4317 4318 4319 | |
add_initial_node_analysis(node_promise, track_history=True)
Adds a data exploration/analysis node based on a node promise.
Automatically captures history for undo/redo support.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_promise
|
NodePromise
|
The promise representing the node to be analyzed. |
required |
track_history
|
bool
|
Whether to track this change in history (default True). |
True
|
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2919 2920 2921 2922 2923 2924 2925 2926 2927 2928 2929 2930 2931 2932 2933 2934 2935 2936 2937 2938 2939 2940 2941 | |
add_join(join_settings)
Adds a join node to combine two data streams based on key columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
join_settings
|
NodeJoin
|
The settings for the join operation. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3335 3336 3337 3338 3339 3340 3341 3342 3343 3344 3345 3346 3347 3348 3349 3350 3351 3352 3353 3354 3355 3356 3357 3358 3359 3360 3361 3362 3363 3364 3365 3366 3367 3368 3369 3370 3371 3372 3373 3374 3375 3376 3377 3378 | |
add_kafka_source(node_kafka_source)
Adds a node to read data from a Kafka or Redpanda topic.
Follows the same pattern as add_database_reader: offloads consumption to the worker, which writes an IPC temp file and returns a serialized LazyFrame reference. Offset tracking is handled by Kafka consumer groups.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_kafka_source
|
NodeKafkaSource
|
The settings for the Kafka source node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4903 4904 4905 4906 4907 4908 4909 4910 4911 4912 4913 4914 4915 4916 4917 4918 4919 4920 4921 4922 4923 4924 4925 4926 4927 4928 4929 4930 4931 4932 4933 4934 4935 4936 4937 4938 4939 4940 4941 4942 4943 4944 4945 4946 4947 4948 4949 4950 4951 4952 4953 4954 4955 4956 4957 4958 4959 4960 4961 4962 4963 4964 4965 4966 4967 4968 4969 4970 4971 4972 4973 4974 4975 4976 4977 4978 4979 4980 4981 4982 4983 4984 4985 4986 4987 4988 4989 4990 4991 4992 4993 4994 4995 4996 4997 4998 4999 5000 5001 5002 5003 5004 5005 5006 5007 5008 5009 5010 5011 5012 5013 5014 5015 5016 5017 5018 5019 5020 5021 5022 5023 5024 5025 5026 5027 5028 5029 5030 5031 5032 5033 5034 5035 5036 5037 5038 5039 5040 5041 5042 5043 5044 5045 5046 5047 5048 5049 5050 5051 5052 5053 5054 5055 5056 5057 5058 5059 5060 5061 5062 5063 5064 | |
add_manual_input(input_file)
Adds a node for manual data entry.
This is a convenience alias for add_datasource.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_file
|
NodeManualInput
|
The settings and data for the manual input node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5600 5601 5602 5603 5604 5605 5606 5607 5608 | |
add_missing_user_defined_node(*, user_defined_node_settings, node_type, error)
Adds a placeholder for a custom node that cannot be loaded on this machine.
The stored settings are preserved verbatim (lossless re-save), the node
renders with its connections, and running the flow fails this node with
error instead of silently dropping it.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 | |
add_node_promise(node_promise, track_history=True)
Adds a placeholder node to the graph that is not yet fully configured.
Useful for building the graph structure before all settings are available. Automatically captures history for undo/redo support.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_promise
|
NodePromise
|
A promise object containing basic node information. |
required |
track_history
|
bool
|
Whether to track this change in history (default True). |
True
|
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 | |
add_node_step(node_id, function, input_columns=None, output_schema=None, node_type=None, drop_columns=None, renew_schema=True, setting_input=None, cache_results=None, schema_callback=None, input_node_ids=None)
The core method for adding or updating a node in the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int | str
|
The unique ID for the node. |
required |
function
|
Callable
|
The core processing function for the node. |
required |
input_columns
|
list[str]
|
A list of input column names required by the function. |
None
|
output_schema
|
list[FlowfileColumn]
|
A predefined schema for the node's output. |
None
|
node_type
|
str
|
A string identifying the type of node (e.g., 'filter', 'join'). |
None
|
drop_columns
|
list[str]
|
A list of columns to be dropped after the function executes. |
None
|
renew_schema
|
bool
|
If True, the schema is recalculated after execution. |
True
|
setting_input
|
Any
|
A configuration object containing settings for the node. |
None
|
cache_results
|
bool
|
If True, the node's results are cached for future runs. |
None
|
schema_callback
|
Callable
|
A function that dynamically calculates the output schema. |
None
|
input_node_ids
|
list[int]
|
A list of IDs for the nodes that this node depends on. |
None
|
Returns:
| Type | Description |
|---|---|
FlowNode
|
The created or updated FlowNode object. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4189 4190 4191 4192 4193 4194 4195 4196 4197 4198 4199 4200 4201 4202 4203 4204 4205 4206 4207 4208 4209 4210 4211 4212 4213 4214 4215 4216 4217 4218 4219 4220 4221 4222 4223 4224 4225 4226 4227 4228 4229 4230 4231 4232 4233 4234 4235 4236 4237 4238 4239 4240 4241 4242 4243 4244 4245 4246 4247 4248 4249 4250 4251 4252 4253 4254 4255 4256 4257 4258 4259 4260 4261 4262 4263 4264 4265 4266 4267 4268 4269 4270 4271 4272 4273 4274 4275 4276 4277 4278 4279 4280 4281 4282 4283 4284 4285 4286 4287 4288 4289 4290 4291 4292 4293 4294 4295 4296 4297 4298 4299 4300 4301 4302 4303 4304 4305 4306 | |
add_node_to_starting_list(node)
Adds a node to the list of starting nodes for the flow if not already present.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node
|
FlowNode
|
The FlowNode to add as a starting node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2136 2137 2138 2139 2140 2141 2142 2143 | |
add_nodes_to_group(group_id, node_ids)
Add nodes to an existing group and refit its bounds.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 | |
add_output(output_file)
Adds an output node to write the final data to a destination.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
output_file
|
NodeOutput
|
The settings for the output file. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4321 4322 4323 4324 4325 4326 4327 4328 4329 4330 4331 4332 4333 4334 4335 4336 4337 4338 4339 4340 4341 4342 4343 4344 4345 4346 4347 4348 4349 4350 4351 4352 4353 4354 4355 4356 4357 4358 4359 4360 4361 4362 4363 4364 4365 4366 4367 4368 4369 4370 | |
add_pivot(pivot_settings)
Adds a pivot node to the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pivot_settings
|
NodePivot
|
The settings for the pivot operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2842 2843 2844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 2861 2862 2863 2864 2865 2866 2867 2868 2869 2870 2871 2872 2873 2874 2875 2876 2877 2878 | |
add_polars_code(node_polars_code)
Adds a node that executes custom Polars code.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_polars_code
|
NodePolarsCode
|
The settings for the Polars code node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3067 3068 3069 3070 3071 3072 3073 3074 3075 3076 3077 3078 3079 3080 3081 3082 3083 3084 3085 3086 3087 3088 3089 3090 | |
add_python_script(node_python_script)
Adds a node that executes Python code on a kernel container.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3129 3130 3131 3132 3133 3134 3135 3136 3137 3138 3139 3140 3141 3142 3143 3144 3145 3146 3147 3148 3149 3150 3151 3152 3153 3154 3155 3156 3157 3158 3159 3160 3161 3162 3163 3164 3165 3166 3167 3168 3169 3170 3171 3172 3173 3174 3175 3176 3177 3178 3179 3180 3181 3182 3183 | |
add_random_split(settings)
Adds a node that randomly partitions rows into N labeled outputs.
Returns a NamedOutputs; the framework unpacks it into
_named_outputs so each split is reachable via its own output handle.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
NodeRandomSplit
|
The settings object specifying the splits and optional seed. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4000 4001 4002 4003 4004 4005 4006 4007 4008 4009 4010 4011 4012 4013 4014 4015 4016 4017 4018 4019 4020 4021 4022 4023 4024 4025 4026 4027 4028 4029 4030 4031 4032 4033 | |
add_read(input_file)
Adds a node to read data from a local file (e.g., CSV, Parquet, Excel).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_file
|
NodeRead
|
The settings for the read operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5469 5470 5471 5472 5473 5474 5475 5476 5477 5478 5479 5480 5481 5482 5483 5484 5485 5486 5487 5488 5489 5490 5491 5492 5493 5494 5495 5496 5497 5498 5499 5500 5501 5502 5503 5504 5505 5506 5507 5508 5509 5510 5511 5512 5513 5514 5515 5516 5517 5518 5519 5520 5521 5522 5523 5524 5525 5526 5527 5528 5529 5530 5531 5532 5533 5534 5535 5536 5537 5538 5539 5540 5541 5542 5543 5544 5545 5546 5547 5548 5549 5550 5551 5552 5553 5554 5555 5556 | |
add_record_count(node_number_of_records)
Adds a filter node to the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_number_of_records
|
NodeRecordCount
|
The settings for the record count operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3048 3049 3050 3051 3052 3053 3054 3055 3056 3057 3058 3059 3060 3061 3062 3063 3064 3065 | |
add_record_id(record_id_settings)
Adds a node to create a new column with a unique ID for each record.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
record_id_settings
|
NodeRecordId
|
The settings object specifying the name of the new record ID column. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4035 4036 4037 4038 4039 4040 4041 4042 4043 4044 4045 4046 4047 4048 4049 4050 4051 4052 4053 4054 4055 4056 4057 | |
add_rest_api_reader(node_rest_api_reader)
Adds a node that reads from a REST API.
All network I/O (HTTP round-trips, pagination, retries) is offloaded to
the worker via ExternalRestApiFetcher — the core never makes the
external call. The credential is resolved to an encrypted token here
(from the user's secret store, or an inline plaintext) and the worker
decrypts it just-in-time. A generic API's columns are unknown until a
response is fetched, so schema_callback returns the columns cached on
the node by the "Fetch sample" action — empty until the user samples or
runs, in which case the fetched frame defines the schema.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5222 5223 5224 5225 5226 5227 5228 5229 5230 5231 5232 5233 5234 5235 5236 5237 5238 5239 5240 5241 5242 5243 5244 5245 5246 5247 5248 5249 5250 5251 5252 5253 5254 5255 5256 5257 5258 5259 5260 5261 5262 5263 5264 5265 5266 5267 5268 5269 5270 5271 5272 5273 5274 5275 5276 5277 5278 5279 5280 5281 5282 5283 5284 5285 5286 5287 5288 5289 5290 5291 5292 5293 5294 5295 5296 5297 5298 5299 5300 5301 | |
add_run_flow(settings)
Adds a node that executes a catalog-registered flow as a subflow.
Inputs are keyed: handle input-0 carries optional parameter data; handles input-1..input-N feed the subflow's flow_input nodes (input_slots order). Outputs mirror the subflow's flow_output nodes (output_slots order).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4439 4440 4441 4442 4443 4444 4445 4446 4447 4448 4449 4450 4451 4452 4453 4454 4455 4456 4457 4458 4459 4460 4461 4462 4463 4464 4465 4466 4467 4468 4469 4470 4471 4472 4473 4474 4475 4476 4477 4478 4479 4480 4481 4482 4483 4484 4485 4486 4487 4488 4489 4490 4491 4492 4493 4494 | |
add_sample(sample_settings)
Adds a node to take a random or top-N sample of the data.
Every method stays lazy, so the node needs no local/remote branch: the sample is part of the plan the worker receives, not a materialised frame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sample_settings
|
NodeSample
|
The settings object specifying the sampling method, the size or fraction to keep, and an optional seed. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3969 3970 3971 3972 3973 3974 3975 3976 3977 3978 3979 3980 3981 3982 3983 3984 3985 3986 3987 3988 3989 3990 3991 3992 3993 3994 3995 3996 3997 3998 | |
add_select(select_settings)
Adds a node to select, rename, reorder, or drop columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
select_settings
|
NodeSelect
|
The settings for the select operation. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4087 4088 4089 4090 4091 4092 4093 4094 4095 4096 4097 4098 4099 4100 4101 4102 4103 4104 4105 4106 4107 4108 4109 4110 4111 4112 4113 4114 4115 4116 4117 4118 4119 4120 4121 4122 4123 4124 4125 4126 4127 4128 4129 | |
add_sort(sort_settings)
Adds a node to sort the data based on one or more columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sort_settings
|
NodeSort
|
The settings for the sort operation. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3946 3947 3948 3949 3950 3951 3952 3953 3954 3955 3956 3957 3958 3959 3960 3961 3962 3963 3964 3965 3966 3967 | |
add_sql_query(node_sql_query)
Adds a node that executes a SQL query against connected data sources.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_sql_query
|
NodeSqlQuery
|
The settings for the SQL query node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3092 3093 3094 3095 3096 3097 3098 3099 3100 3101 3102 3103 3104 3105 3106 3107 3108 3109 3110 3111 3112 3113 3114 3115 3116 3117 3118 3119 3120 3121 3122 3123 3124 3125 3126 3127 | |
add_sql_source(external_source_input)
Adds a node that reads data from a SQL source.
This is a convenience alias for add_external_source.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
external_source_input
|
NodeExternalSource
|
The settings for the external SQL source node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5066 5067 5068 5069 5070 5071 5072 5073 5074 5075 | |
add_text_to_rows(node_text_to_rows)
Adds a node that splits cell values into multiple rows.
This is useful for un-nesting data where a single field contains multiple values separated by a delimiter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_text_to_rows
|
NodeTextToRows
|
The settings object that specifies the column to split and the delimiter to use. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3879 3880 3881 3882 3883 3884 3885 3886 3887 3888 3889 3890 3891 3892 3893 3894 3895 3896 3897 3898 3899 3900 3901 3902 3903 3904 | |
add_train_model(train_settings)
Adds a Train Model node.
Fits a regression model on the worker, stores the serialised artifact
in the global catalog (via :class:ArtifactService), and passes the
input data through unchanged so downstream nodes can keep transforming.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
train_settings
|
NodeTrainModel
|
Settings (model name, target/features, model_type, params). |
required |
Returns:
| Name | Type | Description |
|---|---|---|
The |
FlowGraph
|
class: |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3380 3381 3382 3383 3384 3385 3386 3387 3388 3389 3390 3391 3392 3393 3394 3395 3396 3397 3398 3399 3400 3401 3402 3403 3404 3405 3406 3407 3408 3409 3410 3411 3412 3413 3414 3415 3416 3417 3418 3419 3420 3421 3422 3423 3424 3425 3426 3427 3428 3429 3430 3431 3432 3433 3434 3435 3436 3437 3438 3439 3440 3441 3442 3443 3444 3445 3446 3447 3448 3449 3450 3451 3452 3453 3454 3455 3456 3457 3458 3459 3460 3461 3462 3463 3464 3465 3466 3467 3468 3469 3470 3471 3472 3473 3474 3475 3476 3477 3478 3479 3480 3481 3482 3483 3484 3485 3486 3487 3488 3489 3490 3491 3492 3493 3494 3495 3496 3497 3498 3499 3500 3501 3502 3503 3504 3505 3506 3507 3508 3509 3510 3511 3512 3513 3514 3515 3516 3517 3518 3519 3520 3521 3522 3523 3524 3525 3526 3527 3528 3529 3530 3531 3532 3533 3534 3535 3536 3537 3538 3539 3540 3541 3542 3543 3544 3545 3546 3547 3548 3549 3550 3551 3552 3553 3554 3555 3556 3557 3558 3559 3560 3561 3562 3563 3564 3565 3566 3567 3568 3569 3570 3571 3572 3573 3574 3575 3576 3577 3578 3579 3580 3581 3582 3583 3584 3585 3586 3587 3588 3589 3590 3591 3592 3593 3594 3595 3596 3597 3598 3599 | |
add_union(union_settings)
Adds a union node to combine multiple data streams.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
union_settings
|
NodeUnion
|
The settings for the union operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2899 2900 2901 2902 2903 2904 2905 2906 2907 2908 2909 2910 2911 2912 2913 2914 2915 2916 2917 | |
add_unique(unique_settings)
Adds a node to find and remove duplicate rows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
unique_settings
|
NodeUnique
|
The settings for the unique operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3205 3206 3207 3208 3209 3210 3211 3212 3213 3214 3215 3216 3217 3218 3219 3220 3221 3222 3223 | |
add_unpivot(unpivot_settings)
Adds an unpivot node to the graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
unpivot_settings
|
NodeUnpivot
|
The settings for the unpivot operation. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2880 2881 2882 2883 2884 2885 2886 2887 2888 2889 2890 2891 2892 2893 2894 2895 2896 2897 | |
add_user_defined_node(*, custom_node, user_defined_node_settings)
Adds a user-defined custom node to the graph.
When the custom node has a kernel_id set, the process code is sent
to the kernel for execution instead of running locally. This enables
custom nodes to use external packages installed on the kernel.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
custom_node
|
CustomNodeBase
|
The custom node instance to add. |
required |
user_defined_node_settings
|
UserDefinedNode
|
The settings for the user-defined node. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 2398 2399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 | |
add_wait_for(settings)
Adds a Wait For node — passes the left input through and waits on the right.
Two distinct input handles like Join: connect the data path to the left and the dependency node (e.g. Train Model) to the right. The right input's data is discarded; only its completion gates this node.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3793 3794 3795 3796 3797 3798 3799 3800 3801 3802 3803 3804 3805 3806 3807 3808 3809 3810 3811 3812 3813 3814 3815 3816 3817 3818 3819 3820 3821 3822 3823 | |
add_window_functions(settings)
Adds a window-functions node (rolling, cumulative, rank, tile).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
NodeWindowFunctions
|
The settings for the window-functions operation. |
required |
Returns:
| Type | Description |
|---|---|
FlowGraph
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
3906 3907 3908 3909 3910 3911 3912 3913 3914 3915 3916 3917 3918 3919 3920 3921 3922 3923 3924 3925 3926 3927 3928 3929 3930 3931 3932 3933 3934 3935 3936 3937 3938 3939 3940 3941 3942 3943 3944 | |
apply_layout(y_spacing=150, x_spacing=200, initial_y=100)
Calculates and applies a layered layout to all nodes in the graph.
This updates their x and y positions for UI rendering.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
y_spacing
|
int
|
The minimum vertical spacing between two nodes in a layer. |
150
|
x_spacing
|
int
|
The horizontal spacing between layers. |
200
|
initial_y
|
int
|
The y-position of the topmost node. |
100
|
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 | |
assign_node_to_named_group(node_id, name, *, color=None)
Assign a node to a group identified by name, creating it if absent (find-or-create).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2095 2096 2097 2098 2099 2100 2101 2102 | |
cancel()
Cancels an ongoing graph execution.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6385 6386 6387 6388 6389 6390 6391 6392 | |
capture_history_if_changed(pre_snapshot, action_type, description, node_id=None)
Capture history only if the flow state actually changed.
Use this for settings updates where the change might be a no-op. Call this AFTER the change is applied.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pre_snapshot
|
FlowfileData
|
The FlowfileData captured BEFORE the change. |
required |
action_type
|
HistoryActionType
|
The type of action that was performed. |
required |
description
|
str
|
Human-readable description of the action. |
required |
node_id
|
int
|
Optional ID of the affected node. |
None
|
Returns:
| Type | Description |
|---|---|
bool
|
True if a change was detected and snapshot was captured. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 | |
capture_history_snapshot(action_type, description, node_id=None)
Capture the current state before a change for undo support.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
action_type
|
HistoryActionType
|
The type of action being performed. |
required |
description
|
str
|
Human-readable description of the action. |
required |
node_id
|
int
|
Optional ID of the affected node. |
None
|
Returns:
| Type | Description |
|---|---|
bool
|
True if snapshot was captured, False if skipped. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 | |
check_flow_laziness()
Check whether the flow supports lazy execution for virtual tables.
Finds all catalog-writer nodes in the graph and checks whether their upstream dependencies are fully lazy. Only the nodes that actually feed into a catalog writer matter — unrelated branches (e.g. an Explore Data node on a separate path) are ignored.
Returns a tuple of (is_optimizable, reasons_if_not).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5656 5657 5658 5659 5660 5661 5662 5663 5664 5665 5666 5667 5668 5669 5670 5671 5672 5673 5674 5675 5676 5677 5678 5679 5680 | |
close_flow()
Performs cleanup operations, such as clearing node caches.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6394 6395 6396 6397 6398 | |
copy_node(new_node_settings, existing_setting_input, node_type)
Creates a copy of an existing node.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
new_node_settings
|
NodePromise
|
The promise containing new settings (like ID and position). |
required |
existing_setting_input
|
Any
|
The settings object from the node being copied. |
required |
node_type
|
str
|
The type of the node being copied. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6652 6653 6654 6655 6656 6657 6658 6659 6660 6661 6662 6663 6664 6665 6666 6667 6668 6669 6670 6671 6672 6673 6674 6675 6676 6677 6678 6679 6680 6681 6682 6683 6684 6685 6686 6687 6688 6689 6690 | |
create_group(name, node_ids, *, color=None, bounds=None, parent_group_id=None, child_group_ids=None)
Create a visual group. Organizational only.
Members are the given nodes (group_id) and child groups (their parent_group_id). The new group itself nests under parent_group_id. Bounds are computed when not supplied.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 | |
delete_group(group_id)
Remove a group box (ungroup). Members and sub-groups lift up one level.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 | |
delete_node(node_id)
Deletes a node from the graph and updates all its connections.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int | str
|
The ID of the node to delete. |
required |
Raises:
| Type | Description |
|---|---|
Exception
|
If the node with the given ID does not exist. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
4136 4137 4138 4139 4140 4141 4142 4143 4144 4145 4146 4147 4148 4149 4150 4151 4152 4153 4154 4155 4156 4157 4158 4159 4160 4161 4162 4163 4164 4165 4166 4167 4168 4169 4170 4171 4172 4173 4174 4175 4176 4177 4178 4179 4180 4181 4182 | |
generate_code()
Generates code for the flow graph. This method exports the flow graph to a Polars-compatible format.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6714 6715 6716 6717 6718 6719 6720 | |
get_frontend_data()
Formats the graph structure into a JSON-like dictionary for a specific legacy frontend.
This method transforms the graph's state into a format compatible with the Drawflow.js library.
Returns:
| Type | Description |
|---|---|
dict
|
A dictionary representing the graph in Drawflow format. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6554 6555 6556 6557 6558 6559 6560 6561 6562 6563 6564 6565 6566 6567 6568 6569 6570 6571 6572 6573 6574 6575 6576 6577 6578 6579 6580 6581 6582 6583 6584 6585 6586 6587 6588 6589 6590 6591 6592 6593 6594 6595 6596 6597 6598 6599 6600 6601 6602 6603 6604 6605 6606 6607 6608 6609 6610 6611 6612 6613 6614 6615 6616 6617 6618 6619 6620 6621 6622 6623 6624 6625 6626 | |
get_history_state()
Get the current state of the history system.
Returns:
| Type | Description |
|---|---|
HistoryState
|
HistoryState with information about available undo/redo operations. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1748 1749 1750 1751 1752 1753 1754 | |
get_implicit_starter_nodes()
Finds nodes that can act as starting points but are not explicitly defined as such.
Some nodes, like the Polars Code node, can function without an input. This method identifies such nodes if they have no incoming connections.
Returns:
| Type | Description |
|---|---|
list[FlowNode]
|
A list of |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5687 5688 5689 5690 5691 5692 5693 5694 5695 5696 5697 5698 5699 5700 5701 | |
get_node(node_id=None)
Retrieves a node from the graph by its ID.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int | str
|
The ID of the node to retrieve. If None, retrieves the last added node. |
None
|
Returns:
| Type | Description |
|---|---|
FlowNode | None
|
The FlowNode object, or None if not found. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2344 2345 2346 2347 2348 2349 2350 2351 2352 2353 2354 2355 2356 2357 | |
get_node_data(node_id, include_example=True)
Retrieves all data needed to render a node in the UI.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int
|
The ID of the node. |
required |
include_example
|
bool
|
Whether to include data samples in the result. |
True
|
Returns:
| Type | Description |
|---|---|
NodeData
|
A NodeData object, or None if the node is not found. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6302 6303 6304 6305 6306 6307 6308 6309 6310 6311 6312 6313 | |
get_node_storage()
Serializes the entire graph's state into a storable format.
Returns:
| Type | Description |
|---|---|
FlowInformation
|
A FlowInformation object representing the complete graph. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6366 6367 6368 6369 6370 6371 6372 6373 6374 6375 6376 6377 6378 6379 6380 6381 6382 6383 | |
get_nodes_overview()
Gets a list of dictionary representations for all nodes in the graph.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2328 2329 2330 2331 2332 2333 | |
get_run_info()
Gets a summary of the most recent graph execution.
Returns:
| Type | Description |
|---|---|
RunInformation
|
A RunInformation object with details about the last run. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6266 6267 6268 6269 6270 6271 6272 6273 6274 6275 6276 6277 6278 6279 6280 6281 | |
get_vue_flow_input()
Formats the graph's nodes and edges into a schema suitable for the VueFlow frontend.
Returns:
| Type | Description |
|---|---|
VueFlowInput
|
A VueFlowInput object. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6628 6629 6630 6631 6632 6633 6634 6635 6636 6637 6638 6639 6640 6641 6642 6643 6644 | |
has_unsaved_changes()
Return True if the flow has changed since the last save point.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1760 1761 1762 | |
mark_as_saved()
Mark the current flow state as the saved baseline (for dirty tracking).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1756 1757 1758 | |
print_tree()
Print flow_graph as a visual tree structure, showing the DAG relationships with ASCII art.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 | |
redo()
Redo the last undone action.
Returns:
| Type | Description |
|---|---|
UndoRedoResult
|
UndoRedoResult indicating success or failure. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1740 1741 1742 1743 1744 1745 1746 | |
release_run()
Release the single-run slot claimed by try_claim_run (idempotent).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5772 5773 5774 5775 | |
remove_from_output_cols(columns)
Removes specified columns from the list of expected output columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
columns
|
list[str]
|
A list of column names to remove. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2335 2336 2337 2338 2339 2340 2341 2342 | |
remove_nodes_from_group(node_ids)
Remove nodes from whatever group they belong to; prune groups left empty.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 | |
reset()
Forces a deep reset on all nodes in the graph.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6646 6647 6648 6649 6650 | |
restore_from_snapshot(snapshot)
Clear current state and rebuild from a snapshot.
This method is used internally by undo/redo to restore a previous state.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
snapshot
|
FlowfileData
|
The FlowfileData snapshot to restore from. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 | |
restore_groups(groups)
Replace the runtime group registry (used by open_flow and restore_from_snapshot).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2126 2127 2128 2129 2130 2131 2132 | |
run_graph()
Executes the entire data flow graph from start to finish.
Independent nodes within the same execution stage are run in parallel using threads. Stages are processed sequentially so that all dependencies are satisfied before a stage begins.
Returns:
| Type | Description |
|---|---|
RunInformation | None
|
A RunInformation object summarizing the execution results. |
Raises:
| Type | Description |
|---|---|
Exception
|
If the flow is already running. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6214 6215 6216 6217 6218 6219 6220 6221 6222 6223 6224 6225 6226 6227 6228 6229 6230 6231 6232 6233 6234 6235 6236 6237 6238 6239 6240 6241 6242 6243 6244 6245 6246 6247 6248 6249 6250 6251 6252 6253 6254 6255 6256 6257 6258 6259 6260 6261 6262 6263 6264 | |
save_flow(flow_path)
Saves the current state of the flow graph to a file.
Supports multiple formats based on file extension: - .yaml / .yml: New YAML format - .json: JSON format
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_path
|
str
|
The path where the flow file will be saved. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
6419 6420 6421 6422 6423 6424 6425 6426 6427 6428 6429 6430 6431 6432 6433 6434 6435 6436 6437 6438 6439 6440 6441 6442 6443 6444 6445 6446 6447 6448 6449 6450 6451 6452 6453 6454 6455 6456 6457 6458 6459 6460 6461 6462 6463 6464 6465 6466 6467 6468 | |
set_group_bounds(updates)
Persist group box bounds (used together with set_node_positions on drag/resize).
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2116 2117 2118 2119 2120 2121 2122 2123 2124 | |
set_node_positions(updates)
Persist dragged node positions (absolute canvas coordinates) onto setting_input.
Plain mutator: the caller (update_layout route) captures history once for the whole drag-end batch so node moves and group-bounds changes share one snapshot.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 | |
trigger_fetch_node(node_id, *, performance_mode=False, reset_cache=True)
Executes a specific node in the graph by its ID.
The defaults are the data-preview contract: a non-performance run, so the
node stores its result and can serve the 100-row example grid. Callers
that only need the node's query plan (the Explore Data drawer) pass
performance_mode=True, which skips that store entirely, and
reset_cache=False so exploring doesn't evict a useful cache.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5777 5778 5779 5780 5781 5782 5783 5784 5785 5786 5787 5788 5789 5790 5791 5792 5793 5794 5795 5796 5797 5798 5799 5800 5801 5802 5803 5804 5805 5806 5807 5808 5809 5810 5811 5812 5813 5814 5815 5816 5817 5818 5819 5820 5821 5822 5823 5824 5825 5826 5827 5828 5829 5830 5831 5832 5833 5834 5835 | |
try_claim_run()
Atomically claim the flow's single-run slot; False when a run is already in flight.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
5764 5765 5766 5767 5768 5769 5770 | |
undo()
Undo the last action by restoring to the previous state.
Returns:
| Type | Description |
|---|---|
UndoRedoResult
|
UndoRedoResult indicating success or failure. |
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
1732 1733 1734 1735 1736 1737 1738 | |
update_group(group_id, *, name=None, color=None, bounds=None, collapsed=None)
Rename / recolor / move / resize / collapse a group box.
Source code in flowfile_core/flowfile_core/flowfile/flow_graph.py
2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 | |
FlowNode
The FlowNode represents a single operation in the FlowGraph. Each node corresponds to a specific transformation or action, such as filtering or grouping data.
flowfile_core.flowfile.flow_node.flow_node.FlowNode
Represents a single node in a data flow graph.
This class manages the node's state, its data processing function, and its connections to other nodes within the graph.
Methods:
| Name | Description |
|---|---|
__call__ |
Makes the node instance callable, acting as an alias for execute_node. |
__init__ |
Initializes a FlowNode instance. |
__repr__ |
Provides a string representation of the FlowNode instance. |
add_lead_to_in_depend_source |
Ensures this node is registered in the |
add_node_connection |
Adds a connection from a source node to this node. |
calculate_hash |
Calculates a hash based on settings and input node hashes. |
cancel |
Cancels an ongoing external process if one is running. |
check_upstream_laziness |
Check whether all upstream dependencies of this node support lazy execution. |
clear_table_example |
Clear the table example in the results so that it clears the existing results |
create_schema_callback_from_function |
Wraps a node's function to create a schema callback that extracts the schema. |
delete_input_node |
Removes a connection from a specific input node. |
delete_lead_to_node |
Removes a connection to a specific downstream node. |
evaluate_nodes |
Triggers a state reset for all directly connected downstream nodes. |
execute_full_local |
Backward-compatible alias for _do_execute_full_local. |
execute_local |
Backward-compatible alias for _do_execute_local_with_sampling. |
execute_node |
Execute the node based on its current state and settings. |
execute_remote |
Backward-compatible alias for _do_execute_remote. |
get_all_dependent_node_ids |
Yields the IDs of all downstream nodes recursively. |
get_all_dependent_nodes |
Yields all downstream nodes recursively. |
get_column_stats |
Computes on-demand stats for one column of this node's cached result. |
get_edge_input |
Generates |
get_flow_file_column_schema |
Retrieves the schema for a specific column from the output schema. |
get_input_type |
Gets the type of connection ('main', 'left', 'right') for a given input node ID. |
get_node_data |
Gathers all necessary data for representing the node in the UI. |
get_node_information |
Updates and returns the node's information object. |
get_node_input |
Creates a |
get_output |
Get the result for a specific output handle. |
get_output_data |
Gets the full output data sample for this node. |
get_predicted_resulting_data |
Creates a |
get_predicted_schema |
Predicts the output schema of the node without full execution. |
get_repr |
Gets a detailed dictionary representation of the node's state. |
get_resulting_data |
Executes the node's function to produce the actual output data. |
get_table_example |
Generates a |
invalidate_cache |
Force cache invalidation by incrementing the cache epoch. |
needs_reset |
Checks if the node's hash has changed, indicating an outdated state. |
needs_run |
Determines if the node needs to be executed. |
peek_output_engine |
Passively resolves the cached result engine for an output handle. |
post_init |
Reset every instance attribute to its default state. |
prepare_before_run |
Resets results and errors before a new execution. |
print |
Helper method to log messages with node context. |
remap_dynamic_inputs |
Re-key keyed connections after the node's input slots changed. |
remove_cache |
Removes cached results for this node. |
reset |
Resets the node's execution state and schema information. |
schema_for_handle |
Return the cached schema for a specific output handle. |
set_node_information |
Populates the |
store_example_data_generator |
Stores a generator function for fetching a sample of the result data. |
update_node |
Updates the properties of the node. |
Attributes:
| Name | Type | Description |
|---|---|---|
accepts_dynamic_inputs |
bool
|
True when this node's connections are keyed by target handle (run_flow). |
all_inputs |
list[FlowNode]
|
Gets a list of all nodes connected to any input port. |
executor |
NodeExecutor
|
Lazy-initialized executor instance. |
function |
Callable
|
Gets the core processing function of the node. |
has_input |
bool
|
Checks if this node has any input connections. |
has_next_step |
bool
|
Checks if this node has any downstream connections. |
hash |
str
|
Gets the cached hash for the node, calculating it if it doesn't exist. |
is_correct |
bool
|
Checks if the node's input connections satisfy its template requirements. |
is_setup |
bool
|
Checks if the node has been properly configured and is ready for execution. |
is_start |
bool
|
Determines if the node is a starting node in the flow. |
left_input |
Optional[FlowNode]
|
Gets the node connected to the left input port. |
main_input |
list[FlowNode]
|
Gets the list of nodes connected to the main input port(s). |
name |
str
|
Gets the name of the node. |
node_id |
str | int
|
Gets the unique identifier of the node. |
number_of_leads_to_nodes |
int | None
|
Counts the number of downstream node connections. |
right_input |
Optional[FlowNode]
|
Gets the node connected to the right input port. |
schema |
list[FlowfileColumn]
|
Gets the definitive output schema of the node. |
schema_callback |
SingleExecutionFuture
|
Gets the schema callback function, creating one if it doesn't exist. |
setting_input |
Any
|
Gets the node's specific configuration settings. |
singular_input |
bool
|
Checks if the node template specifies exactly one input. |
singular_main_input |
FlowNode
|
Gets the input node, assuming it is a single-input type. |
state_needs_reset |
bool
|
Checks if the node's state needs to be reset. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 | |
accepts_dynamic_inputs
property
True when this node's connections are keyed by target handle (run_flow).
all_inputs
property
Gets a list of all nodes connected to any input port.
Returns:
| Type | Description |
|---|---|
list[FlowNode]
|
A list of all input FlowNodes. |
executor
property
Lazy-initialized executor instance.
Reusing the same executor avoids object creation overhead when execute_node is called multiple times.
function
property
writable
Gets the core processing function of the node.
Returns:
| Type | Description |
|---|---|
Callable
|
The callable function. |
has_input
property
Checks if this node has any input connections.
Returns:
| Type | Description |
|---|---|
bool
|
True if it has at least one input node. |
has_next_step
property
Checks if this node has any downstream connections.
Returns:
| Type | Description |
|---|---|
bool
|
True if it has at least one downstream node. |
hash
property
Gets the cached hash for the node, calculating it if it doesn't exist.
Returns:
| Type | Description |
|---|---|
str
|
The string hash value. |
is_correct
property
Checks if the node's input connections satisfy its template requirements.
Returns:
| Type | Description |
|---|---|
bool
|
True if connections are valid, False otherwise. |
is_setup
property
Checks if the node has been properly configured and is ready for execution.
Returns:
| Type | Description |
|---|---|
bool
|
True if the node is set up, False otherwise. |
is_start
property
Determines if the node is a starting node in the flow.
A starting node requires no inputs.
Returns:
| Type | Description |
|---|---|
bool
|
True if the node is a start node, False otherwise. |
left_input
property
Gets the node connected to the left input port.
Returns:
| Type | Description |
|---|---|
Optional[FlowNode]
|
The left input FlowNode, or None. |
main_input
property
Gets the list of nodes connected to the main input port(s).
Returns:
| Type | Description |
|---|---|
list[FlowNode]
|
A list of main input FlowNodes. |
name
property
writable
Gets the name of the node.
Returns:
| Type | Description |
|---|---|
str
|
The node's name. |
node_id
property
Gets the unique identifier of the node.
Returns:
| Type | Description |
|---|---|
str | int
|
The node's ID. |
number_of_leads_to_nodes
property
Counts the number of downstream node connections.
Returns:
| Type | Description |
|---|---|
int | None
|
The number of nodes this node leads to. |
right_input
property
Gets the node connected to the right input port.
Returns:
| Type | Description |
|---|---|
Optional[FlowNode]
|
The right input FlowNode, or None. |
schema
property
Gets the definitive output schema of the node.
If not already run, it falls back to the predicted schema.
Returns:
| Type | Description |
|---|---|
list[FlowfileColumn]
|
A list of FlowfileColumn objects. |
schema_callback
property
writable
Gets the schema callback function, creating one if it doesn't exist.
The callback is used for predicting the output schema without full execution.
Returns:
| Type | Description |
|---|---|
SingleExecutionFuture
|
A SingleExecutionFuture instance wrapping the schema function. |
setting_input
property
writable
Gets the node's specific configuration settings.
Returns:
| Type | Description |
|---|---|
Any
|
The settings object. |
singular_input
property
Checks if the node template specifies exactly one input.
Returns:
| Type | Description |
|---|---|
bool
|
True if the node is a single-input type. |
singular_main_input
property
Gets the input node, assuming it is a single-input type.
Returns:
| Type | Description |
|---|---|
FlowNode
|
The single input FlowNode, or None. |
state_needs_reset
property
writable
Checks if the node's state needs to be reset.
Returns:
| Type | Description |
|---|---|
bool
|
True if a reset is required, False otherwise. |
__call__(*args, **kwargs)
Makes the node instance callable, acting as an alias for execute_node.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1375 1376 1377 | |
__init__(node_id, function, parent_uuid, setting_input, name, node_type, input_columns=None, output_schema=None, drop_columns=None, renew_schema=True, pos_x=0, pos_y=0, schema_callback=None)
Initializes a FlowNode instance.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
str | int
|
Unique identifier for the node. |
required |
function
|
Callable
|
The core data processing function for the node. |
required |
parent_uuid
|
str
|
The UUID of the parent flow. |
required |
setting_input
|
Any
|
The configuration/settings object for the node. |
required |
name
|
str
|
The name of the node. |
required |
node_type
|
str
|
The type identifier of the node (e.g., 'join', 'filter'). |
required |
input_columns
|
list[str]
|
List of column names expected as input. |
None
|
output_schema
|
list[FlowfileColumn]
|
The schema of the columns to be added. |
None
|
drop_columns
|
list[str]
|
List of column names to be dropped. |
None
|
renew_schema
|
bool
|
Flag to indicate if the schema should be renewed. |
True
|
pos_x
|
float
|
The x-coordinate on the canvas. |
0
|
pos_y
|
float
|
The y-coordinate on the canvas. |
0
|
schema_callback
|
Callable
|
A custom function to calculate the output schema. |
None
|
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 | |
__repr__()
Provides a string representation of the FlowNode instance.
Returns:
| Type | Description |
|---|---|
str
|
A string showing the node's ID and type. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1836 1837 1838 1839 1840 1841 1842 | |
add_lead_to_in_depend_source()
Ensures this node is registered in the leads_to_nodes list of its inputs.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1279 1280 1281 1282 1283 | |
add_node_connection(from_node, insert_type='main', output_handle=DEFAULT_OUTPUT_HANDLE, target_handle=None)
Adds a connection from a source node to this node.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
from_node
|
FlowNode
|
The node to connect from. |
required |
insert_type
|
Literal['main', 'left', 'right']
|
The type of input to connect to ('main', 'left', 'right'). |
'main'
|
output_handle
|
str
|
The output handle on the source node (e.g. 'output-0', 'output-1'). |
DEFAULT_OUTPUT_HANDLE
|
target_handle
|
str | None
|
For dynamic-input nodes only: the target handle the edge lands on ('input-0'..'input-N'). Ignored for static nodes. |
None
|
Raises:
| Type | Description |
|---|---|
Exception
|
If the insert_type is invalid. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 | |
calculate_hash(setting_input)
Calculates a hash based on settings and input node hashes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
setting_input
|
Any
|
The node's settings object to be included in the hash. |
required |
Returns:
| Type | Description |
|---|---|
str
|
A string hash value. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 | |
cancel()
Cancels an ongoing external process if one is running.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 | |
check_upstream_laziness()
Check whether all upstream dependencies of this node support lazy execution.
Walks the DAG backwards from this node (excluding itself) and reports any eager or conditional nodes that would prevent a lazy/optimized execution path.
Returns:
| Type | Description |
|---|---|
bool
|
A tuple of (is_lazy, reasons). |
list[str]
|
upstream node has |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 | |
clear_table_example()
Clear the table example in the results so that it clears the existing results Returns: None
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 | |
create_schema_callback_from_function(f)
Wraps a node's function to create a schema callback that extracts the schema.
For multi-output functions, every handle's schema is captured in
_named_schemas on the single call; the callback itself still
returns the default handle's schema so the existing contract holds.
Thread-safe: uses _execution_lock to prevent concurrent execution with get_resulting_data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
f
|
Callable
|
The node's core function that returns a FlowDataEngine or NamedOutputs. |
required |
Returns:
| Type | Description |
|---|---|
Callable[[], list[FlowfileColumn]]
|
A callable that, when executed, returns the default output's schema. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 | |
delete_input_node(node_id, connection_type='input-0', complete=False)
Removes a connection from a specific input node.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int
|
The ID of the input node to disconnect. |
required |
connection_type
|
InputConnectionClass
|
The specific input handle (e.g., 'input-0', 'input-1'). |
'input-0'
|
complete
|
bool
|
If True, tries to delete from all input types. |
False
|
Returns:
| Type | Description |
|---|---|
bool
|
True if a connection was found and removed, False otherwise. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 | |
delete_lead_to_node(node_id)
Removes a connection to a specific downstream node.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int
|
The ID of the downstream node to disconnect. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if the connection was found and removed, False otherwise. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 | |
evaluate_nodes(deep=False)
Triggers a state reset for all directly connected downstream nodes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
deep
|
bool
|
If True, the reset propagates recursively through the entire downstream graph. |
False
|
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
861 862 863 864 865 866 867 868 869 | |
execute_full_local(performance_mode=False)
Backward-compatible alias for _do_execute_full_local.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1617 1618 1619 | |
execute_local(flow_id, performance_mode=False)
Backward-compatible alias for _do_execute_local_with_sampling.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1621 1622 1623 | |
execute_node(run_location, reset_cache=False, performance_mode=False, retry=True, node_logger=None, optimize_for_downstream=True)
Execute the node based on its current state and settings.
Delegates all execution and skip logic to the NodeExecutor, which is the single source of truth for deciding whether a node should run.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
run_location
|
ExecutionLocationsLiteral
|
Where to execute ('local' or 'remote') |
required |
reset_cache
|
bool
|
Force cache invalidation |
False
|
performance_mode
|
bool
|
Skip example data generation for speed |
False
|
retry
|
bool
|
Allow retry on recoverable errors |
True
|
node_logger
|
NodeLogger | None
|
Logger for this node's execution |
None
|
optimize_for_downstream
|
bool
|
Cache wide transforms for downstream nodes |
True
|
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 | |
execute_remote(performance_mode=False, node_logger=None)
Backward-compatible alias for _do_execute_remote.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1625 1626 1627 | |
get_all_dependent_node_ids()
Yields the IDs of all downstream nodes recursively.
Returns:
| Type | Description |
|---|---|
None
|
A generator of all dependent node IDs. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1295 1296 1297 1298 1299 1300 1301 1302 1303 | |
get_all_dependent_nodes()
Yields all downstream nodes recursively.
Returns:
| Type | Description |
|---|---|
None
|
A generator of all dependent FlowNode objects. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1285 1286 1287 1288 1289 1290 1291 1292 1293 | |
get_column_stats(column_name, output_handle=DEFAULT_OUTPUT_HANDLE, offload_to_worker=False)
Computes on-demand stats for one column of this node's cached result.
The stats land on the result engine's FlowfileColumn (the single
source of truth), so this run's later previews carry them too; the
return value is that column's ordinary FileColumn representation.
offload_to_worker ships the aggregate to the worker — used for
locally-run flows, whose cached engine is the full upstream plan.
Raises ColumnStatsUnavailable when there is no cached result to
aggregate over without executing or re-pulling data, and
pl.exceptions.ColumnNotFoundError for an unknown column.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 | |
get_edge_input()
Generates NodeEdge objects for all input connections to this node.
Returns:
| Type | Description |
|---|---|
list[NodeEdge]
|
A list of |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 | |
get_flow_file_column_schema(col_name)
Retrieves the schema for a specific column from the output schema.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
col_name
|
str
|
The name of the column. |
required |
Returns:
| Type | Description |
|---|---|
FlowfileColumn | None
|
The FlowfileColumn object for that column, or None if not found. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
871 872 873 874 875 876 877 878 879 880 881 882 | |
get_input_type(node_id)
Gets the type of connection ('main', 'left', 'right') for a given input node ID.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id
|
int
|
The ID of the input node. |
required |
Returns:
| Type | Description |
|---|---|
list
|
A list of connection types for that node ID. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 | |
get_node_data(flow_id, include_example=False, include_output=True, include_inputs=True)
Gathers all necessary data for representing the node in the UI.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the parent flow. |
required |
include_example
|
bool
|
If True, includes data samples. |
False
|
include_output
|
bool
|
If True, computes this node's own output preview
( |
True
|
include_inputs
|
bool
|
If True, resolves each connected input's schema
( |
True
|
Returns:
| Type | Description |
|---|---|
NodeData
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 | |
get_node_information()
Updates and returns the node's information object.
Returns:
| Type | Description |
|---|---|
NodeInformation
|
The |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
624 625 626 627 628 629 630 631 | |
get_node_input()
Creates a NodeInput schema object for representing this node in the UI.
Returns:
| Type | Description |
|---|---|
NodeInput
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 | |
get_output(handle=DEFAULT_OUTPUT_HANDLE)
Get the result for a specific output handle.
For nodes with multiple outputs (e.g. kernel-based custom nodes),
returns the FlowDataEngine associated with the given handle.
Falls back to the default results.resulting_data for single-output nodes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
handle
|
str
|
The output handle identifier (e.g. |
DEFAULT_OUTPUT_HANDLE
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine | None
|
The FlowDataEngine for the requested output, or None. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 | |
get_output_data()
Gets the full output data sample for this node.
Returns:
| Type | Description |
|---|---|
TableExample
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
2157 2158 2159 2160 2161 2162 2163 | |
get_predicted_resulting_data(handle=DEFAULT_OUTPUT_HANDLE)
Creates a FlowDataEngine instance based on the predicted schema.
This avoids executing the node's full logic. For multi-output nodes the
handle argument selects which output's schema to reflect so that a
downstream node wired to e.g. output-1 sees that partition's schema.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
handle
|
str
|
The output handle to reflect. Ignored for single-output nodes. |
DEFAULT_OUTPUT_HANDLE
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A FlowDataEngine instance with a schema but no data. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 | |
get_predicted_schema(force=False)
Predicts the output schema of the node without full execution.
It uses the schema_callback or infers from predicted data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
force
|
bool
|
If True, forces recalculation even if a predicted schema exists. |
False
|
Returns:
| Type | Description |
|---|---|
list[FlowfileColumn] | None
|
A list of FlowfileColumn objects representing the predicted schema. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 | |
get_repr()
Gets a detailed dictionary representation of the node's state.
Returns:
| Type | Description |
|---|---|
dict
|
A dictionary containing key information about the node. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 | |
get_resulting_data()
Executes the node's function to produce the actual output data.
Handles both regular functions and external data sources.
Thread-safe and single-flight: the node's own _execution_lock ensures
the function runs at most once and the result is memoized, so N downstream
consumers materialize it once. A node acquires only its OWN lock; upstream
inputs are read through each upstream's own get_resulting_data(), so
lock acquisition always follows the DAG (a node -> its parents) and cannot
form a cross-node cycle.
Returns:
| Type | Description |
|---|---|
FlowDataEngine | None
|
A FlowDataEngine instance containing the result, or None on error. |
Raises:
| Type | Description |
|---|---|
Exception
|
Propagates exceptions from the node's function execution. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 | |
get_table_example(include_data=False, output_handle=DEFAULT_OUTPUT_HANDLE)
Generates a TableExample model summarizing the node's output.
This can optionally include a sample of the data. For multi-output
nodes, output_handle selects which named output to preview.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
include_data
|
bool
|
If True, includes a data sample in the result. |
False
|
output_handle
|
str
|
The output handle to preview (e.g. |
DEFAULT_OUTPUT_HANDLE
|
Returns:
| Type | Description |
|---|---|
TableExample | None
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 | |
invalidate_cache()
Force cache invalidation by incrementing the cache epoch.
Changes the node's hash so Development mode re-executes instead of returning stale results. Used after external state changes (e.g. Kafka consumer group offset reset) that don't alter the node's configuration.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 | |
needs_reset()
Checks if the node's hash has changed, indicating an outdated state.
Returns:
| Type | Description |
|---|---|
bool
|
True if the calculated hash differs from the stored hash. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1716 1717 1718 1719 1720 1721 1722 | |
needs_run(performance_mode, node_logger=None, execution_location='remote')
Determines if the node needs to be executed.
The decision is based on its run state, caching settings, and execution mode.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
performance_mode
|
bool
|
True if the flow is in performance mode. |
required |
node_logger
|
NodeLogger
|
The logger instance for this node. |
None
|
execution_location
|
ExecutionLocationsLiteral
|
The target execution location. |
'remote'
|
Returns:
| Type | Description |
|---|---|
bool
|
True if the node should be run, False otherwise. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 | |
peek_output_engine(output_handle=DEFAULT_OUTPUT_HANDLE)
Passively resolves the cached result engine for an output handle.
Never routes through get_output()/get_resulting_data(), which
can re-execute the node. Identity checks only: FlowDataEngine truthiness
goes through __len__, which can trigger a full collect.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 | |
post_init()
Reset every instance attribute to its default state.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 | |
prepare_before_run()
Resets results and errors before a new execution.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1629 1630 1631 1632 1633 1634 1635 1636 | |
print(v)
Helper method to log messages with node context.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
v
|
Any
|
The message or value to log. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
986 987 988 989 990 991 992 | |
remap_dynamic_inputs(mapping)
Re-key keyed connections after the node's input slots changed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
mapping
|
dict[str, str | None]
|
old handle -> new handle, or None to drop that connection. Handles absent from the mapping keep their key. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, list[str]]
|
{"moved": [...], "dropped": [...]} describing what happened, so the |
dict[str, list[str]]
|
API layer can surface removed connections to the UI. |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 | |
remove_cache()
Removes cached results for this node.
Note: Currently not fully implemented.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1330 1331 1332 1333 1334 1335 1336 1337 1338 | |
reset(deep=False)
Resets the node's execution state and schema information.
This also triggers a reset on all downstream nodes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
deep
|
bool
|
If True, forces a reset even if the hash hasn't changed. |
False
|
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 | |
schema_for_handle(handle)
Return the cached schema for a specific output handle.
Falls back to the default schema property when the handle is unknown
or the node is single-output, so callers can always rely on this.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
288 289 290 291 292 293 294 295 296 297 298 | |
set_node_information()
Populates the node_information attribute with the current state.
This includes the node's connections, settings, and position.
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 | |
store_example_data_generator(external_df_fetcher)
Stores a generator function for fetching a sample of the result data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
external_df_fetcher
|
ExternalDfFetcher | ExternalSampler
|
The process that generated the sample data. |
required |
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 | |
update_node(function, input_columns=None, output_schema=None, drop_columns=None, name=None, setting_input=None, pos_x=0, pos_y=0, schema_callback=None)
Updates the properties of the node.
This is called during initialization and when settings are changed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
function
|
Callable
|
The new core data processing function. |
required |
input_columns
|
list[str]
|
The new list of input columns. |
None
|
output_schema
|
list[FlowfileColumn]
|
The new schema of added columns. |
None
|
drop_columns
|
list[str]
|
The new list of dropped columns. |
None
|
name
|
str
|
The new name for the node. |
None
|
setting_input
|
Any
|
The new settings object. |
None
|
pos_x
|
float
|
The new x-coordinate. |
0
|
pos_y
|
float
|
The new y-coordinate. |
0
|
schema_callback
|
Callable
|
The new custom schema callback function. |
None
|
Source code in flowfile_core/flowfile_core/flowfile/flow_node/flow_node.py
415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 | |
The FlowDataEngine
The FlowDataEngine is the primary engine of the library, providing a rich API for data manipulation, I/O, and transformation. Its methods are grouped below by functionality.
flowfile_core.flowfile.flow_data_engine.flow_data_engine.FlowDataEngine
dataclass
The core data handling engine for Flowfile.
This class acts as a high-level wrapper around a Polars DataFrame or LazyFrame, providing a unified API for data ingestion, transformation, and output. It manages data state (lazy vs. eager), schema information, and execution logic.
Attributes:
| Name | Type | Description |
|---|---|---|
_data_frame |
DataFrame | LazyFrame
|
The underlying Polars DataFrame or LazyFrame. |
columns |
list[Any]
|
A list of column names in the current data frame. |
name |
str
|
An optional name for the data engine instance. |
number_of_records |
int
|
The number of records. Can be -1 for lazy frames. |
errors |
list
|
A list of errors encountered during operations. |
_schema |
list[FlowfileColumn] | None
|
A cached list of |
Methods:
| Name | Description |
|---|---|
__call__ |
Makes the class instance callable, returning itself. |
__get_sample__ |
Internal method to get a sample of the data. |
__getitem__ |
Accesses a specific column or item from the DataFrame. |
__init__ |
Initializes the FlowDataEngine from various data sources. |
__len__ |
Returns the number of records in the table. |
__repr__ |
Returns a string representation of the FlowDataEngine. |
add_new_values |
Adds a new column with the provided values. |
add_record_id |
Adds a record ID (row number) column to the DataFrame. |
align_to_schema |
Aligns the DataFrame to an expected schema. |
apply_dynamic_rename |
Renames a subset of columns according to a single rule. |
apply_flowfile_formula |
Applies a formula to create a new column or transform an existing one. |
apply_sql_formula |
Applies an SQL-style formula using |
assert_equal |
Asserts that this DataFrame is equal to another. |
cache |
Caches the current DataFrame to disk and updates the internal reference. |
calculate_schema |
Calculates and returns the schema. |
change_column_types |
Changes the data type of one or more columns. |
collect |
Collects the data and returns it as a Polars DataFrame. |
collect_external |
Materializes data from a tracked external source. |
concat |
Concatenates this DataFrame with one or more other DataFrames. |
count |
Gets the total number of records. |
create_from_external_source |
Creates a FlowDataEngine from an external data source. |
create_from_path |
Creates a FlowDataEngine from a local file path. |
create_from_path_worker |
Creates a FlowDataEngine from a path in a worker process. |
create_from_schema |
Creates an empty FlowDataEngine from a schema definition. |
create_from_sql |
Creates a FlowDataEngine by executing a SQL query. |
create_random |
Creates a FlowDataEngine with randomly generated data. |
do_cross_join |
Performs a cross join with another DataFrame. |
do_filter |
Filters rows based on a predicate expression. |
do_group_by |
Performs a group-by operation on the DataFrame. |
do_pivot |
Converts the DataFrame from a long to a wide format, aggregating values. |
do_select |
Performs a complex column selection, renaming, and reordering operation. |
do_sort |
Sorts the DataFrame by one or more columns. |
do_window_functions |
Applies window functions (rolling, cumulative, rank, tile) to the data. |
drop_columns |
Drops specified columns from the DataFrame. |
filter_split |
Partition rows by |
from_cloud_storage_obj |
Creates a FlowDataEngine from an object in cloud storage. |
generate_enumerator |
Generates a FlowDataEngine with a single column containing a sequence of integers. |
get_estimated_file_size |
Estimates the file size in bytes if the data originated from a local file. |
get_number_of_records |
Gets the total number of records in the DataFrame. |
get_number_of_records_in_process |
Get the number of records in the DataFrame in the local process. |
get_output_sample |
Gets a sample of the data as a list of dictionaries. |
get_record_count |
Returns a new FlowDataEngine with a single column 'number_of_records' |
get_sample |
Gets a sample of rows from the DataFrame. |
get_schema_column |
Retrieves the schema information for a single column by its name. |
get_select_inputs |
Gets |
get_subset |
Gets the first |
initialize_empty_fl |
Initializes an empty LazyFrame. |
iter_batches |
Iterates over the DataFrame in batches. |
join |
Performs a standard SQL-style join with another DataFrame. |
known_record_count |
Returns the exact record count only when it is already known for free. |
make_unique |
Gets the unique rows from the DataFrame. |
output |
Writes the DataFrame to a local output file. |
random_sample |
Takes a uniform random sample of rows without materialising the frame. |
random_split |
Randomly partition rows into N labeled groups (in-process). |
random_split_external |
Worker-offloaded variant of :meth: |
reorganize_order |
Reorganizes columns into a specified order. |
resolve_dynamic_rename_map |
Compute the |
save |
Saves the DataFrame to a file in a separate thread. |
select_columns |
Selects a subset of columns from the DataFrame. |
set_streamable |
Sets whether DataFrame operations should be streamable. |
shallow_copy |
Cheap de-aliasing wrapper around the same (immutable) Polars frame. |
solve_graph |
Solves a graph problem represented by 'from' and 'to' columns. |
split |
Splits a column's text values into multiple rows based on a delimiter. |
start_fuzzy_join |
Starts a fuzzy join operation in a background process. |
to_arrow |
Converts the DataFrame to a PyArrow Table. |
to_cloud_storage_obj |
Writes the DataFrame to an object in cloud storage. |
to_database_obj |
Writes the DataFrame to a SQL database in-process (local execution path). |
to_dict |
Converts the DataFrame to a Python dictionary of columns. |
to_pylist |
Converts the DataFrame to a list of Python dictionaries. |
to_raw_data |
Converts the DataFrame to a |
unpivot |
Converts the DataFrame from a wide to a long format. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 1860 1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 1874 1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 1894 1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 1960 1961 1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 2202 2203 2204 2205 2206 2207 2208 2209 2210 2211 2212 2213 2214 2215 2216 2217 2218 2219 2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 2231 2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 2252 2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 2327 2328 2329 2330 2331 2332 2333 2334 2335 2336 2337 2338 2339 2340 2341 2342 2343 2344 2345 2346 2347 2348 2349 2350 2351 2352 2353 2354 2355 2356 2357 2358 2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 2398 2399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 2443 2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 2465 2466 2467 2468 2469 2470 2471 2472 2473 2474 2475 2476 2477 2478 2479 2480 2481 2482 2483 2484 2485 2486 2487 2488 2489 2490 2491 2492 2493 2494 2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 2579 2580 2581 2582 2583 2584 2585 2586 2587 2588 2589 2590 2591 2592 2593 2594 2595 2596 2597 2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 2627 2628 2629 2630 2631 2632 2633 2634 2635 2636 2637 2638 2639 2640 2641 2642 2643 2644 2645 2646 2647 2648 2649 2650 2651 2652 2653 2654 2655 2656 2657 2658 2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 2691 2692 2693 2694 2695 2696 2697 2698 2699 2700 2701 2702 2703 2704 2705 2706 2707 2708 2709 2710 2711 2712 2713 2714 2715 2716 2717 2718 2719 2720 2721 2722 2723 2724 2725 2726 2727 2728 2729 2730 2731 2732 2733 2734 2735 2736 2737 2738 2739 2740 2741 2742 2743 2744 2745 2746 2747 2748 2749 2750 2751 2752 2753 2754 2755 2756 2757 2758 2759 2760 2761 2762 2763 2764 2765 2766 2767 2768 2769 2770 2771 2772 2773 2774 2775 2776 2777 2778 2779 2780 2781 2782 2783 2784 2785 2786 2787 2788 2789 2790 2791 2792 2793 2794 2795 2796 2797 2798 2799 2800 2801 2802 2803 2804 2805 2806 2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 2819 2820 2821 2822 2823 2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 2843 2844 2845 2846 2847 2848 2849 2850 2851 2852 2853 2854 2855 2856 2857 2858 2859 2860 2861 2862 2863 2864 2865 2866 2867 2868 2869 2870 2871 | |
__name__
property
The name of the table.
cols_idx
property
A dictionary mapping column names to their integer index.
data_frame
property
writable
The underlying Polars DataFrame or LazyFrame.
This property provides access to the Polars object that backs the FlowDataEngine. It handles lazy-loading from external sources if necessary.
Returns:
| Type | Description |
|---|---|
LazyFrame | DataFrame | None
|
The active Polars |
external_source
property
The external data source, if any.
has_errors
property
Checks if there are any errors.
lazy
property
writable
Indicates if the DataFrame is in lazy mode.
number_of_fields
property
The number of columns (fields) in the DataFrame.
Returns:
| Type | Description |
|---|---|
int
|
The integer count of columns. |
schema
property
The schema of the DataFrame as a list of FlowfileColumn objects.
This property lazily calculates the schema if it hasn't been determined yet.
Returns:
| Type | Description |
|---|---|
list[FlowfileColumn]
|
A list of |
__call__()
Makes the class instance callable, returning itself.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1596 1597 1598 | |
__get_sample__(n_rows=100, streamable=True)
Internal method to get a sample of the data.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 | |
__getitem__(item)
Accesses a specific column or item from the DataFrame.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
832 833 834 | |
__init__(raw_data=None, path_ref=None, name=None, optimize_memory=True, schema=None, number_of_records=None, calculate_schema_stats=False, streamable=True, number_of_records_callback=None, data_callback=None)
Initializes the FlowDataEngine from various data sources.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
raw_data
|
list[dict] | list[Any] | dict[str, Any] | ParquetFile | DataFrame | LazyFrame | RawData
|
The input data. Can be a list of dicts, a Polars DataFrame/LazyFrame,
or a |
None
|
path_ref
|
str
|
A string path to a Parquet file. |
None
|
name
|
str
|
An optional name for the data engine instance. |
None
|
optimize_memory
|
bool
|
If True, prefers lazy operations to conserve memory. |
True
|
schema
|
list[FlowfileColumn] | list[str] | Schema
|
An optional schema definition. Can be a list of |
None
|
number_of_records
|
int
|
The number of records, if known. |
None
|
calculate_schema_stats
|
bool
|
If True, computes detailed statistics for each column. |
False
|
streamable
|
bool
|
If True, allows for streaming operations when possible. |
True
|
number_of_records_callback
|
Callable
|
A callback function to retrieve the number of records. |
None
|
data_callback
|
Callable
|
A callback function to retrieve the data. |
None
|
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 | |
__len__()
Returns the number of records in the table.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1600 1601 1602 | |
__repr__()
Returns a string representation of the FlowDataEngine.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1592 1593 1594 | |
add_new_values(values, col_name=None)
Adds a new column with the provided values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
values
|
Iterable
|
An iterable (e.g., list, tuple) of values to add as a new column. |
required |
col_name
|
str
|
The name for the new column. Defaults to 'new_values'. |
None
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 | |
add_record_id(record_id_settings)
Adds a record ID (row number) column to the DataFrame.
Can generate a simple sequential ID or a grouped ID that resets for each group.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
record_id_settings
|
RecordIdInput
|
A |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 | |
align_to_schema(expected_schema)
Aligns the DataFrame to an expected schema.
Adds any missing columns as typed nulls and reorders all columns to match the order defined in expected_schema. Extra columns present in the data but absent from the expected schema are appended at the end so no data is silently dropped.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expected_schema
|
list[FlowfileColumn]
|
The desired column list, in order. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2394 2395 2396 2397 2398 2399 2400 2401 2402 2403 2404 2405 2406 2407 2408 2409 2410 2411 2412 2413 2414 2415 2416 2417 2418 2419 2420 2421 2422 2423 2424 2425 2426 2427 2428 2429 | |
apply_dynamic_rename(settings)
Renames a subset of columns according to a single rule.
Supports prefix, suffix, flowfile-formula, and first-row rename modes, with
column selection by name list, by data type, or across all columns. In
"first_row" mode the first row is always dropped from the output after
its values are promoted to column headers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
DynamicRenameInput
|
The dynamic rename configuration. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
FlowDataEngine
|
unchanged if the rule resolves to no renames). |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2659 2660 2661 2662 2663 2664 2665 2666 2667 2668 2669 2670 2671 2672 2673 2674 2675 2676 2677 2678 2679 2680 2681 2682 2683 2684 2685 2686 2687 2688 2689 2690 | |
apply_flowfile_formula(func, col_name, output_data_type=None)
Applies a formula to create a new column or transform an existing one.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
func
|
str
|
A string containing a Polars expression formula. |
required |
col_name
|
str
|
The name of the new or transformed column. |
required |
output_data_type
|
DataType
|
The desired Polars data type for the output column. |
None
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2444 2445 2446 2447 2448 2449 2450 2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 | |
apply_sql_formula(func, col_name, output_data_type=None)
Applies an SQL-style formula using pl.sql_expr.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
func
|
str
|
A string containing an SQL expression. |
required |
col_name
|
str
|
The name of the new or transformed column. |
required |
output_data_type
|
DataType
|
The desired Polars data type for the output column. |
None
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2692 2693 2694 2695 2696 2697 2698 2699 2700 2701 2702 2703 2704 2705 2706 2707 2708 2709 | |
assert_equal(other, ordered=True, strict_schema=False)
Asserts that this DataFrame is equal to another.
Useful for testing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
other
|
FlowDataEngine
|
The other |
required |
ordered
|
bool
|
If True, the row order must be identical. |
True
|
strict_schema
|
bool
|
If True, the data types of the schemas must be identical. |
False
|
Raises:
| Type | Description |
|---|---|
Exception
|
If the DataFrames are not equal based on the specified criteria. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 2184 2185 2186 2187 2188 2189 2190 2191 2192 2193 2194 2195 2196 2197 2198 2199 2200 2201 | |
cache()
Caches the current DataFrame to disk and updates the internal reference.
This triggers a background process to write the current LazyFrame's result
to a temporary file. Subsequent operations on this FlowDataEngine instance
will read from the cached file, which can speed up downstream computations.
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
The same |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 | |
calculate_schema()
Calculates and returns the schema.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2854 2855 2856 2857 | |
change_column_types(transforms, calculate_schema=False)
Changes the data type of one or more columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
transforms
|
list[SelectInput]
|
A list of |
required |
calculate_schema
|
bool
|
If True, recalculates the schema after the type change. |
False
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 | |
collect(n_records=None)
Collects the data and returns it as a Polars DataFrame.
This method triggers the execution of the lazy query plan (if applicable) and returns the result. It supports streaming to optimize memory usage for large datasets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_records
|
int
|
The maximum number of records to collect. If None, all records are collected. |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
A Polars |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 | |
collect_external()
Materializes data from a tracked external source.
If the FlowDataEngine was created from an ExternalDataSource, this
method will trigger the data retrieval, update the internal _data_frame
to a LazyFrame of the collected data, and reset the schema to be
re-evaluated.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 | |
concat(other)
Concatenates this DataFrame with one or more other DataFrames.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
other
|
Iterable[FlowDataEngine] | FlowDataEngine
|
A single |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2758 2759 2760 2761 2762 2763 2764 2765 2766 2767 2768 2769 2770 2771 | |
count()
Gets the total number of records.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2859 2860 2861 | |
create_from_external_source(external_source)
classmethod
Creates a FlowDataEngine from an external data source.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
external_source
|
ExternalDataSource
|
An object that conforms to the |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 | |
create_from_path(received_table)
classmethod
Creates a FlowDataEngine from a local file path.
Supports various file types like CSV, Parquet, and Excel.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
received_table
|
ReceivedTable
|
A |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 | |
create_from_path_worker(received_table, flow_id, node_id)
classmethod
Creates a FlowDataEngine from a path in a worker process.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2863 2864 2865 2866 2867 2868 2869 2870 2871 | |
create_from_schema(schema)
classmethod
Creates an empty FlowDataEngine from a schema definition.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
schema
|
list[FlowfileColumn]
|
A list of |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new, empty |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 | |
create_from_sql(sql, conn)
classmethod
Creates a FlowDataEngine by executing a SQL query.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sql
|
str
|
The SQL query string to execute. |
required |
conn
|
Any
|
A database connection object or connection URI string. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 | |
create_random(number_of_records=1000)
classmethod
Creates a FlowDataEngine with randomly generated data.
Useful for testing and examples.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
number_of_records
|
int
|
The number of random records to generate. |
1000
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 | |
do_cross_join(cross_join_input, auto_generate_selection, verify_integrity, other)
Performs a cross join with another DataFrame.
A cross join produces the Cartesian product of the two DataFrames.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cross_join_input
|
CrossJoinInput
|
A |
required |
auto_generate_selection
|
bool
|
If True, automatically renames columns to avoid conflicts. |
required |
verify_integrity
|
bool
|
If True, checks if the resulting join would be too large. |
required |
other
|
FlowDataEngine
|
The right |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Raises:
| Type | Description |
|---|---|
Exception
|
If |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 | |
do_filter(predicate)
Filters rows based on a predicate expression.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
predicate
|
str
|
A string containing a Polars expression that evaluates to a boolean value. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
FlowDataEngine
|
the predicate. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 | |
do_group_by(group_by_input, calculate_schema_stats=True)
Performs a group-by operation on the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
group_by_input
|
GroupByInput
|
A |
required |
calculate_schema_stats
|
bool
|
If True, calculates schema statistics for the resulting DataFrame. |
True
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 | |
do_pivot(pivot_input, node_logger=None)
Converts the DataFrame from a long to a wide format, aggregating values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pivot_input
|
PivotInput
|
A |
required |
node_logger
|
NodeLogger
|
An optional logger for reporting warnings, e.g., if the pivot column has too many unique values. |
None
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new, pivoted |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 | |
do_select(select_inputs, keep_missing=True)
Performs a complex column selection, renaming, and reordering operation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
select_inputs
|
SelectInputs
|
A |
required |
keep_missing
|
bool
|
If True, columns not specified in |
True
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2773 2774 2775 2776 2777 2778 2779 2780 2781 2782 2783 2784 2785 2786 2787 2788 2789 2790 2791 2792 2793 2794 2795 2796 2797 2798 2799 2800 2801 2802 2803 2804 2805 2806 2807 2808 2809 2810 2811 2812 2813 2814 2815 2816 2817 2818 | |
do_sort(sorts)
Sorts the DataFrame by one or more columns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sorts
|
list[SortByInput]
|
A list of |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 | |
do_window_functions(settings, calculate_schema_stats=False)
Applies window functions (rolling, cumulative, rank, tile) to the data.
When settings.order_by is provided, rows are sorted first so that
rolling and tile operations have a deterministic order; the sort is
preserved in the output. Partitioning (partition_by) is applied via
.over(...) so operations reset for each group.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 | |
drop_columns(columns)
Drops specified columns from the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
columns
|
list[str]
|
A list of column names to drop. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2377 2378 2379 2380 2381 2382 2383 2384 2385 2386 2387 2388 2389 2390 2391 2392 | |
filter_split(predicate)
Partition rows by predicate into pass and fail streams.
Rows where the predicate evaluates to null are dropped from both streams — matching the behaviour of two manually-wired filter nodes with opposing predicates.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 | |
from_cloud_storage_obj(settings)
classmethod
Creates a FlowDataEngine from an object in cloud storage.
This method supports reading from various cloud storage providers like AWS S3, Azure Data Lake Storage, and Google Cloud Storage, with support for various authentication methods.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
CloudStorageReadSettingsInternal
|
A |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the storage type or file format is not supported. |
NotImplementedError
|
If a requested file format like "delta" or "iceberg" is not yet implemented. |
Exception
|
If reading from cloud storage fails. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 | |
generate_enumerator(length=1000, output_name='output_column')
classmethod
Generates a FlowDataEngine with a single column containing a sequence of integers.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
length
|
int
|
The number of integers to generate in the sequence. |
1000
|
output_name
|
str
|
The name of the output column. |
'output_column'
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 | |
get_estimated_file_size()
Estimates the file size in bytes if the data originated from a local file.
This relies on the original path being tracked during file ingestion.
Returns:
| Type | Description |
|---|---|
int
|
The file size in bytes, or 0 if the original path is unknown. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 | |
get_number_of_records(warn=False, force_calculate=False, calculate_in_worker_process=False)
Gets the total number of records in the DataFrame.
For lazy frames, this may trigger a full data scan, which can be expensive.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
warn
|
bool
|
If True, logs a warning if a potentially expensive calculation is triggered. |
False
|
force_calculate
|
bool
|
If True, forces recalculation even if a value is cached. |
False
|
calculate_in_worker_process
|
bool
|
If True, offloads the calculation to a worker process. |
False
|
Returns:
| Type | Description |
|---|---|
int
|
The total number of records. |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the number of records could not be determined. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2253 2254 2255 2256 2257 2258 2259 2260 2261 2262 2263 2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 2275 2276 2277 2278 2279 2280 2281 2282 2283 2284 2285 2286 2287 2288 2289 2290 2291 2292 2293 2294 | |
get_number_of_records_in_process(force_calculate=False)
Get the number of records in the DataFrame in the local process.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
force_calculate
|
bool
|
If True, forces recalculation even if a value is cached. |
False
|
Returns:
| Type | Description |
|---|---|
|
The total number of records. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2220 2221 2222 2223 2224 2225 2226 2227 2228 2229 2230 | |
get_output_sample(n_rows=10)
Gets a sample of the data as a list of dictionaries.
This is typically used to display a preview of the data in a UI.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_rows
|
int
|
The number of rows to sample. |
10
|
Returns:
| Type | Description |
|---|---|
list[dict]
|
A list of dictionaries, where each dictionary represents a row. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 | |
get_record_count()
Returns a new FlowDataEngine with a single column 'number_of_records' containing the total number of records.
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2155 2156 2157 2158 2159 2160 2161 2162 | |
get_sample(n_rows=100, random=False, shuffle=False, seed=None, execution_location=None)
Gets a sample of rows from the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_rows
|
int
|
The number of rows to sample. |
100
|
random
|
bool
|
If True, performs random sampling. If False, takes the first n_rows. |
False
|
shuffle
|
bool
|
If True (and |
False
|
seed
|
int
|
A random seed for reproducibility. |
None
|
execution_location
|
ExecutionLocationsLiteral | None
|
Location which is used to calculate the size of the dataframe |
None
|
Returns:
A new FlowDataEngine instance containing the sampled data.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 | |
get_schema_column(col_name)
Retrieves the schema information for a single column by its name.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
col_name
|
str
|
The name of the column to retrieve. |
required |
Returns:
| Type | Description |
|---|---|
FlowfileColumn
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 | |
get_select_inputs()
Gets SelectInput specifications for all columns in the current schema.
Returns:
| Type | Description |
|---|---|
SelectInputs
|
A |
SelectInputs
|
transformation operations. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2343 2344 2345 2346 2347 2348 2349 2350 2351 2352 | |
get_subset(n_rows=100)
Gets the first n_rows from the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_rows
|
int
|
The number of rows to include in the subset. |
100
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1861 1862 1863 1864 1865 1866 1867 1868 1869 1870 1871 1872 1873 | |
initialize_empty_fl()
Initializes an empty LazyFrame.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2203 2204 2205 2206 2207 | |
iter_batches(batch_size=1000, columns=None)
Iterates over the DataFrame in batches.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
batch_size
|
int
|
The size of each batch. |
1000
|
columns
|
list | tuple | str
|
A list of column names to include in the batches. If None, all columns are included. |
None
|
Yields:
| Type | Description |
|---|---|
FlowDataEngine
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1875 1876 1877 1878 1879 1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 | |
join(join_input, auto_generate_selection, verify_integrity, other)
Performs a standard SQL-style join with another DataFrame.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 2112 2113 2114 2115 2116 2117 2118 2119 2120 | |
known_record_count()
Returns the exact record count only when it is already known for free.
Sources, in order: a previously stored number_of_records (e.g. the
count the worker sent along with a remote run result), or the height of
an eager frame. Returns None otherwise — deliberately never falls back
to get_number_of_records(), which on a lazy frame collects the whole
plan to count it. Stored placeholders are not counts: the cloud readers
stamp CLOUD_PLACEHOLDER_RECORD_COUNT and external-source engines
carry a schema-time 0, so both report unknown here.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2232 2233 2234 2235 2236 2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 | |
make_unique(unique_input=None)
Gets the unique rows from the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
unique_input
|
UniqueInput
|
A |
None
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2744 2745 2746 2747 2748 2749 2750 2751 2752 2753 2754 2755 2756 | |
output(output_fs, flow_id, node_id, execute_remote=False)
Writes the DataFrame to a local output file.
For remote-worker writes the caller (add_output._func) uses
ExternalOutputWriter directly so the fetcher can be exposed on
the node for cancellation; this method only handles the local path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
output_fs
|
OutputSettings
|
An |
required |
flow_id
|
int
|
The flow ID for tracking. |
required |
node_id
|
int | str
|
The node ID for tracking. |
required |
execute_remote
|
bool
|
Retained for signature compatibility; ignored. |
False
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
The same |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2711 2712 2713 2714 2715 2716 2717 2718 2719 2720 2721 2722 2723 2724 2725 2726 2727 2728 2729 2730 2731 2732 2733 2734 2735 2736 2737 2738 2739 2740 2741 2742 | |
random_sample(n=None, fraction=None, seed=None)
Takes a uniform random sample of rows without materialising the frame.
Polars exposes sample only on eager DataFrames, so the lazy
equivalent is built from a shuffled row rank: each row draws a distinct
rank from a random permutation of 0..len, and keeping the ranks
below a threshold keeps a uniform subset. Nothing is collected and the
row count is never queried, so the result stays a plan that ships to
the worker like any other lazy transform — unlike :meth:random_split,
which has to materialise because its outputs must share one permutation.
Sampling more rows than the frame holds yields the whole frame, and the original row order is preserved.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n
|
int | None
|
Number of rows to keep. Mutually exclusive with |
None
|
fraction
|
float | None
|
Share of rows to keep, between 0 and 1. Mutually exclusive with |
None
|
seed
|
int | None
|
Seed for a reproducible sample; None draws a fresh permutation on every execution. |
None
|
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 | |
random_split(splits, seed=None)
Randomly partition rows into N labeled groups (in-process).
Used by add_random_split when execution_location == "local"
(WASM / no-worker). For remote mode the worker-offloaded variant
:meth:random_split_external is used instead.
The shuffled frame is materialized once so that each output shares the
same shuffle — otherwise every handle's .collect() would re-run the
full shuffle+sort independently (O(N·n log n) instead of O(n log n)).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splits
|
list[tuple[str, float]]
|
Ordered (name, percentage) pairs; percentages must sum to
100 (validated upstream in |
required |
seed
|
int | None
|
Random seed; if None, one is generated per call. |
None
|
Returns:
| Type | Description |
|---|---|
NamedOutputs
|
|
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 | |
random_split_external(splits, seed=None, flow_id=-1, node_id=-1)
Worker-offloaded variant of :meth:random_split.
The shuffled frame is materialised once on flowfile_worker (never
in this process). Each returned split is a lazy slice over the
cached parquet, so downstream .collect() on a handle reads only
that split's rows from disk.
Used by add_random_split when execution_location != "local";
the in-process path is :meth:random_split.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 1819 1820 1821 1822 1823 1824 1825 1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 1855 1856 1857 1858 1859 | |
reorganize_order(column_order)
Reorganizes columns into a specified order.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
column_order
|
list[str]
|
A list of column names in the desired order. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2431 2432 2433 2434 2435 2436 2437 2438 2439 2440 2441 2442 | |
resolve_dynamic_rename_map(columns, settings, first_row_values=None)
staticmethod
Compute the {old_name: new_name} map for a dynamic-rename operation.
Pure function — takes the incoming schema as (name, data_type_group) tuples
(where data_type_group is FlowfileColumn.data_type_group, e.g. "Numeric",
"String", "Date", …) and the user's settings, and returns the rename map.
Raises ValueError if the rule would produce duplicate column names.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
columns
|
list[tuple[str, str]]
|
Incoming schema as |
required |
settings
|
DynamicRenameInput
|
The dynamic rename configuration. |
required |
first_row_values
|
dict[str, Any] | None
|
First-row values keyed by original column name. Required
for |
None
|
Returns:
| Type | Description |
|---|---|
dict[str, str]
|
A dict mapping original column name to new column name. No-op renames are |
dict[str, str]
|
omitted, so the result is safe to pass directly to |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2598 2599 2600 2601 2602 2603 2604 2605 2606 2607 2608 2609 2610 2611 2612 2613 2614 2615 2616 2617 2618 2619 2620 2621 2622 2623 2624 2625 2626 | |
save(path, data_type='parquet')
Saves the DataFrame to a file in a separate thread.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str
|
The file path to save to. |
required |
data_type
|
str
|
The format to save in (e.g., 'parquet', 'csv'). |
'parquet'
|
Returns:
| Type | Description |
|---|---|
Future
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 | |
select_columns(list_select)
Selects a subset of columns from the DataFrame.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
list_select
|
list[str] | tuple[str] | str
|
A list, tuple, or single string of column names to select. |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2354 2355 2356 2357 2358 2359 2360 2361 2362 2363 2364 2365 2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 | |
set_streamable(streamable=False)
Sets whether DataFrame operations should be streamable.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2820 2821 2822 | |
shallow_copy()
Cheap de-aliasing wrapper around the same (immutable) Polars frame.
Shares the frame and the cached schema, but owns its own mutable flags (_lazy, _streamable, number_of_records, _schema), so a consumer handed this copy can never mutate an engine shared with sibling consumers. Collect-free: forwarding number_of_records and the cached schema skips both pl.len() and collect_schema() in init (the schema fallback only fires when _schema is unset, and is metadata-only). Deliberately does not carry external_source: memoized results are materialized before they are shared, so the plain frame is the whole result.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2824 2825 2826 2827 2828 2829 2830 2831 2832 2833 2834 2835 2836 2837 2838 2839 2840 2841 2842 2843 2844 2845 | |
solve_graph(graph_solver_input)
Solves a graph problem represented by 'from' and 'to' columns.
This is used for operations like finding connected components in a graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
graph_solver_input
|
GraphSolverInput
|
A |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 | |
split(split_input)
Splits a column's text values into multiple rows based on a delimiter.
This operation is often referred to as "exploding" the DataFrame, as it increases the number of rows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
split_input
|
TextToRowsInput
|
A |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 | |
start_fuzzy_join(fuzzy_match_input, other, file_ref, flow_id=-1, node_id=-1)
Starts a fuzzy join operation in a background process.
This method prepares the data and initiates the fuzzy matching in a separate process, returning a tracker object immediately.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
fuzzy_match_input
|
FuzzyMatchInput
|
A |
required |
other
|
FlowDataEngine
|
The right |
required |
file_ref
|
str
|
A reference string for temporary files. |
required |
flow_id
|
int
|
The flow ID for tracking. |
-1
|
node_id
|
int | str
|
The node ID for tracking. |
-1
|
Returns:
| Type | Description |
|---|---|
ExternalFuzzyMatchFetcher
|
An |
ExternalFuzzyMatchFetcher
|
progress and retrieve the result of the fuzzy join. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1895 1896 1897 1898 1899 1900 1901 1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 1929 1930 1931 1932 | |
to_arrow()
Converts the DataFrame to a PyArrow Table.
This method triggers a .collect() call if the data is lazy,
then converts the resulting eager DataFrame into a pyarrow.Table.
Returns:
| Type | Description |
|---|---|
Table
|
A |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 | |
to_cloud_storage_obj(settings)
Writes the DataFrame to an object in cloud storage.
This method supports writing to various cloud storage providers like AWS S3, Azure Data Lake Storage, and Google Cloud Storage.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
settings
|
CloudStorageWriteSettingsInternal
|
A |
required |
Raises:
| Type | Description |
|---|---|
ValueError
|
If the specified file format is not supported for writing. |
NotImplementedError
|
If the 'append' write mode is used with an unsupported format. |
Exception
|
If the write operation to cloud storage fails for any reason. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 | |
to_database_obj(*, database_type, uri, table_name, if_exists)
Writes the DataFrame to a SQL database in-process (local execution path).
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
465 466 467 468 469 470 471 472 473 474 | |
to_dict()
Converts the DataFrame to a Python dictionary of columns.
Each key in the dictionary is a column name, and the corresponding value is a list of the data in that column.
Returns:
| Type | Description |
|---|---|
dict[str, list]
|
A dictionary mapping column names to lists of their values. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 | |
to_pylist()
Converts the DataFrame to a list of Python dictionaries.
Returns:
| Type | Description |
|---|---|
list[dict]
|
A list where each item is a dictionary representing a row. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1144 1145 1146 1147 1148 1149 1150 1151 1152 | |
to_raw_data()
Converts the DataFrame to a RawData schema object.
Returns:
| Type | Description |
|---|---|
RawData
|
An |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1168 1169 1170 1171 1172 1173 1174 1175 1176 | |
unpivot(unpivot_input)
Converts the DataFrame from a wide to a long format.
This is the inverse of a pivot operation, taking columns and transforming
them into variable and value rows.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
unpivot_input
|
UnpivotInput
|
An |
required |
Returns:
| Type | Description |
|---|---|
FlowDataEngine
|
A new, unpivoted |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_data_engine.py
1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 1396 1397 1398 1399 1400 | |
FlowfileColumn
The FlowfileColumn holds the schema and metadata for a single column managed by the FlowDataEngine.
flowfile_core.flowfile.flow_data_engine.flow_file_column.main.FlowfileColumn
dataclass
Methods:
| Name | Description |
|---|---|
__repr__ |
Provides a concise, developer-friendly representation of the object. |
__str__ |
Provides a detailed, readable summary of the column's metadata. |
apply_statistics |
Applies exactly-computed statistics onto this column. |
Attributes:
| Name | Type | Description |
|---|---|---|
stats_applied |
bool
|
True once exact stats were written via apply_statistics. |
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_file_column/main.py
16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 | |
stats_applied
property
True once exact stats were written via apply_statistics.
__repr__()
Provides a concise, developer-friendly representation of the object. Ideal for debugging and console inspection.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_file_column/main.py
55 56 57 58 59 60 61 62 63 64 65 | |
__str__()
Provides a detailed, readable summary of the column's metadata. It conditionally omits any attribute that is None, ensuring a clean output.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_file_column/main.py
67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 | |
apply_statistics(total_rows, null_count, n_unique=None, min_value=None, max_value=None, average_value=None)
Applies exactly-computed statistics onto this column.
Replaces the schema-time sentinels with real values (from the on-demand column_stats pass) and resets the lazily-derived caches (is_unique, perc_unique, …) so they re-derive from the new counts. Every stat field is overwritten: a value not in this batch (count-only retry, an unsupported dtype) becomes unknown again instead of surviving from an earlier state. Values are stringified and bounded so a long-text min/max can't blow up payloads.
Source code in flowfile_core/flowfile_core/flowfile/flow_data_engine/flow_file_column/main.py
176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 | |
Data Modeling (Schemas)
This section documents the Pydantic models that define the structure of settings and data.
schemas
flowfile_core.schemas.schemas
Classes:
| Name | Description |
|---|---|
CreateGroupRequest |
Body for POST /editor/create_group/. Bounds are optional; computed from members if omitted. |
FlowGraphConfig |
Configuration model for a flow graph's basic properties. |
FlowInformation |
Represents the complete state of a flow, including settings, nodes, and connections. |
FlowSettings |
Extends FlowGraphConfig with additional operational settings for a flow. |
FlowSettingsResponse |
FlowSettings plus runtime-only fields for API responses. Not persisted. |
FlowfileData |
Root model for flowfile serialization (YAML/JSON). |
FlowfileGroup |
Serialized representation of a visual node group (YAML/JSON). |
FlowfileInputConnection |
One keyed input edge of a dynamic-input node (per-edge target handle). |
FlowfileNode |
Node representation for flowfile serialization (YAML/JSON). |
FlowfileSettings |
Settings for flowfile serialization (YAML/JSON). |
GroupBounds |
Axis-aligned bounds of a group box, in absolute canvas coordinates. |
GroupBoundsUpdate |
A single group's new absolute bounds. |
GroupInformation |
Runtime representation of a visual node group (stored in FlowGraph._groups). |
GroupMembershipRequest |
Body for adding/removing nodes from a group. |
NodeConnection |
Represents a connection between two nodes in the flow. |
NodeDefault |
Defines default properties for a node type. |
NodeEdge |
Represents a connection (edge) between two nodes in the frontend. |
NodeInformation |
Stores the state and configuration of a specific node instance within a flow. |
NodeInput |
Represents a node as it is received from the frontend, including position. |
NodePositionUpdate |
A single node's new absolute canvas position. |
NodeTag |
Controlled vocabulary of palette search keywords. |
NodeTemplate |
Defines the template for a node type, specifying its UI and functional characteristics. |
RawLogInput |
Schema for a raw log message. |
UpdateGroupRequest |
Body for POST /editor/update_group/. All fields optional -> partial update. |
UpdateLayoutRequest |
Batch persistence of dragged node positions and/or group bounds (one drag-end -> one call). |
VueFlowInput |
Represents the complete graph structure from the Vue-based frontend. |
Functions:
| Name | Description |
|---|---|
get_global_execution_location |
Calculates the default execution location based on the global settings |
get_settings_class_for_node_type |
Get the settings class for a node type, supporting both standard and user-defined nodes. |
CreateGroupRequest
pydantic-model
Bases: BaseModel
Body for POST /editor/create_group/. Bounds are optional; computed from members if omitted.
Show JSON schema:
{
"description": "Body for POST /editor/create_group/. Bounds are optional; computed from members if omitted.",
"properties": {
"node_ids": {
"items": {
"type": "integer"
},
"title": "Node Ids",
"type": "array"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "X Position"
},
"y_position": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Y Position"
},
"width": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Width"
},
"height": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Height"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
},
"child_group_ids": {
"items": {
"type": "integer"
},
"title": "Child Group Ids",
"type": "array"
}
},
"required": [
"node_ids"
],
"title": "CreateGroupRequest",
"type": "object"
}
Fields:
-
node_ids(list[int]) -
name(str) -
color(GroupColor | None) -
x_position(float | None) -
y_position(float | None) -
width(float | None) -
height(float | None) -
parent_group_id(int | None) -
child_group_ids(list[int])
Source code in flowfile_core/flowfile_core/schemas/schemas.py
807 808 809 810 811 812 813 814 815 816 817 818 | |
FlowGraphConfig
pydantic-model
Bases: BaseModel
Configuration model for a flow graph's basic properties.
Attributes:
| Name | Type | Description |
|---|---|---|
flow_id |
int
|
Unique identifier for the flow. |
description |
Optional[str]
|
A description of the flow. |
save_location |
Optional[str]
|
The location where the flow is saved. |
name |
str
|
The name of the flow. |
path |
str
|
The file path associated with the flow. |
execution_mode |
ExecutionModeLiteral
|
The mode of execution ('Development' or 'Performance'). |
execution_location |
ExecutionLocationsLiteral
|
The location for execution ('local', 'remote'). |
max_parallel_workers |
int
|
Maximum number of threads used for parallel node execution within a stage. Set to 1 to disable parallelism. Defaults to 4. |
parameters |
list[FlowParameter]
|
Flow-level parameters referenceable via ${name} syntax. |
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
}
},
"description": "Configuration model for a flow graph's basic properties.\n\nAttributes:\n flow_id (int): Unique identifier for the flow.\n description (Optional[str]): A description of the flow.\n save_location (Optional[str]): The location where the flow is saved.\n name (str): The name of the flow.\n path (str): The file path associated with the flow.\n execution_mode (ExecutionModeLiteral): The mode of execution ('Development' or 'Performance').\n execution_location (ExecutionLocationsLiteral): The location for execution ('local', 'remote').\n max_parallel_workers (int): Maximum number of threads used for parallel node execution within a\n stage. Set to 1 to disable parallelism. Defaults to 4.\n parameters (list[FlowParameter]): Flow-level parameters referenceable via ${name} syntax.",
"properties": {
"flow_id": {
"description": "Unique identifier for the flow.",
"title": "Flow Id",
"type": "integer"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"save_location": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Save Location"
},
"name": {
"default": "",
"title": "Name",
"type": "string"
},
"path": {
"default": "",
"title": "Path",
"type": "string"
},
"source_registration_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Catalog registration ID when running a registered flow.",
"title": "Source Registration Id"
},
"execution_mode": {
"default": "Performance",
"enum": [
"Development",
"Performance"
],
"title": "Execution Mode",
"type": "string"
},
"execution_location": {
"enum": [
"local",
"remote"
],
"title": "Execution Location",
"type": "string"
},
"max_parallel_workers": {
"default": 4,
"description": "Max threads for parallel node execution.",
"minimum": 1,
"title": "Max Parallel Workers",
"type": "integer"
},
"parameters": {
"description": "Flow-level parameters.",
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameters",
"type": "array"
}
},
"title": "FlowGraphConfig",
"type": "object"
}
Fields:
-
flow_id(int) -
description(str | None) -
save_location(str | None) -
name(str) -
path(str) -
source_registration_id(int | None) -
execution_mode(ExecutionModeLiteral) -
execution_location(ExecutionLocationsLiteral) -
max_parallel_workers(int) -
parameters(list[FlowParameter])
Validators:
-
validate_execution_mode→execution_mode -
validate_and_set_execution_location→execution_location
Source code in flowfile_core/flowfile_core/schemas/schemas.py
139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | |
flow_id
pydantic-field
Unique identifier for the flow.
max_parallel_workers = 4
pydantic-field
Max threads for parallel node execution.
parameters
pydantic-field
Flow-level parameters.
source_registration_id = None
pydantic-field
Catalog registration ID when running a registered flow.
validate_and_set_execution_location(v)
pydantic-validator
Validates and sets the execution location.
1. If None is provided: It defaults to the location determined by global settings.
2. If a value is provided: It checks if the value is compatible with the global
settings. If not (e.g., requesting 'remote' when only 'local' is possible),
it corrects the value to a compatible one.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | |
FlowInformation
pydantic-model
Bases: BaseModel
Represents the complete state of a flow, including settings, nodes, and connections.
Attributes:
| Name | Type | Description |
|---|---|---|
flow_id |
int
|
The unique ID of the flow. |
flow_name |
Optional[str]
|
The name of the flow. |
flow_settings |
FlowSettings
|
The settings for the flow. |
data |
Dict[int, NodeInformation]
|
A dictionary mapping node IDs to their information. |
node_starts |
List[int]
|
A list of starting node IDs. |
node_connections |
List[Tuple[int, int]]
|
A list of tuples representing connections between nodes. |
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
},
"FlowSettings": {
"description": "Extends FlowGraphConfig with additional operational settings for a flow.\n\nAttributes:\n auto_save (bool): Flag to enable or disable automatic saving.\n modified_on (Optional[float]): Timestamp of the last modification.\n show_detailed_progress (bool): Flag to show detailed progress during execution.\n show_edge_labels (bool): Flag to show or hide named edge labels on connections.\n is_running (bool): Indicates if the flow is currently running.\n is_canceled (bool): Indicates if the flow execution has been canceled.\n track_history (bool): Flag to enable or disable undo/redo history tracking.\n validate_settings (bool): Flag to warn on nodes whose settings reference missing input columns.",
"properties": {
"flow_id": {
"description": "Unique identifier for the flow.",
"title": "Flow Id",
"type": "integer"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"save_location": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Save Location"
},
"name": {
"default": "",
"title": "Name",
"type": "string"
},
"path": {
"default": "",
"title": "Path",
"type": "string"
},
"source_registration_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Catalog registration ID when running a registered flow.",
"title": "Source Registration Id"
},
"execution_mode": {
"default": "Performance",
"enum": [
"Development",
"Performance"
],
"title": "Execution Mode",
"type": "string"
},
"execution_location": {
"enum": [
"local",
"remote"
],
"title": "Execution Location",
"type": "string"
},
"max_parallel_workers": {
"default": 4,
"description": "Max threads for parallel node execution.",
"minimum": 1,
"title": "Max Parallel Workers",
"type": "integer"
},
"parameters": {
"description": "Flow-level parameters.",
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameters",
"type": "array"
},
"auto_save": {
"default": false,
"title": "Auto Save",
"type": "boolean"
},
"modified_on": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Modified On"
},
"show_detailed_progress": {
"default": true,
"title": "Show Detailed Progress",
"type": "boolean"
},
"show_edge_labels": {
"default": false,
"title": "Show Edge Labels",
"type": "boolean"
},
"is_running": {
"default": false,
"title": "Is Running",
"type": "boolean"
},
"is_canceled": {
"default": false,
"title": "Is Canceled",
"type": "boolean"
},
"track_history": {
"default": true,
"title": "Track History",
"type": "boolean"
},
"validate_settings": {
"default": true,
"title": "Validate Settings",
"type": "boolean"
}
},
"title": "FlowSettings",
"type": "object"
},
"FlowfileInputConnection": {
"description": "One keyed input edge of a dynamic-input node (per-edge target handle).\n\nOnly nodes whose template sets ``dynamic_inputs`` serialize these; the same\nupstream node may legitimately appear twice with different handles.",
"properties": {
"from_id": {
"title": "From Id",
"type": "integer"
},
"input_handle": {
"title": "Input Handle",
"type": "string"
},
"source_handle": {
"default": "output-0",
"title": "Source Handle",
"type": "string"
}
},
"required": [
"from_id",
"input_handle"
],
"title": "FlowfileInputConnection",
"type": "object"
},
"GroupInformation": {
"description": "Runtime representation of a visual node group (stored in FlowGraph._groups).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"default": 0.0,
"title": "X Position",
"type": "number"
},
"y_position": {
"default": 0.0,
"title": "Y Position",
"type": "number"
},
"width": {
"default": 400.0,
"title": "Width",
"type": "number"
},
"height": {
"default": 250.0,
"title": "Height",
"type": "number"
},
"collapsed": {
"default": false,
"title": "Collapsed",
"type": "boolean"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
}
},
"required": [
"id"
],
"title": "GroupInformation",
"type": "object"
},
"NodeInformation": {
"description": "Stores the state and configuration of a specific node instance within a flow.",
"properties": {
"id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Id"
},
"type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Type"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Is Setup"
},
"is_start_node": {
"default": false,
"title": "Is Start Node",
"type": "boolean"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"x_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "X Position"
},
"y_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "Y Position"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"left_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Input Id"
},
"right_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Input Id"
},
"input_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Input Ids"
},
"outputs": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Outputs"
},
"output_handles": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Handles"
},
"input_connections": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/FlowfileInputConnection"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Connections"
},
"setting_input": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Setting Input"
}
},
"title": "NodeInformation",
"type": "object"
}
},
"description": "Represents the complete state of a flow, including settings, nodes, and connections.\n\nAttributes:\n flow_id (int): The unique ID of the flow.\n flow_name (Optional[str]): The name of the flow.\n flow_settings (FlowSettings): The settings for the flow.\n data (Dict[int, NodeInformation]): A dictionary mapping node IDs to their information.\n node_starts (List[int]): A list of starting node IDs.\n node_connections (List[Tuple[int, int]]): A list of tuples representing connections between nodes.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"flow_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Flow Name"
},
"flow_settings": {
"$ref": "#/$defs/FlowSettings"
},
"data": {
"additionalProperties": {
"$ref": "#/$defs/NodeInformation"
},
"default": {},
"title": "Data",
"type": "object"
},
"node_starts": {
"items": {
"type": "integer"
},
"title": "Node Starts",
"type": "array"
},
"node_connections": {
"default": [],
"items": {
"maxItems": 2,
"minItems": 2,
"prefixItems": [
{
"type": "integer"
},
{
"type": "integer"
}
],
"type": "array"
},
"title": "Node Connections",
"type": "array"
},
"groups": {
"items": {
"$ref": "#/$defs/GroupInformation"
},
"title": "Groups",
"type": "array"
}
},
"required": [
"flow_id",
"flow_settings",
"node_starts"
],
"title": "FlowInformation",
"type": "object"
}
Fields:
-
flow_id(int) -
flow_name(str | None) -
flow_settings(FlowSettings) -
data(dict[int, NodeInformation]) -
node_starts(list[int]) -
node_connections(list[tuple[int, int]]) -
groups(list[GroupInformation])
Validators:
-
ensure_string→flow_name
Source code in flowfile_core/flowfile_core/schemas/schemas.py
699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 | |
ensure_string(v)
pydantic-validator
Validator to ensure the flow_name is always a string. :param v: The value to validate. :return: The value as a string, or an empty string if it's None.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
720 721 722 723 724 725 726 727 | |
FlowSettings
pydantic-model
Bases: FlowGraphConfig
Extends FlowGraphConfig with additional operational settings for a flow.
Attributes:
| Name | Type | Description |
|---|---|---|
auto_save |
bool
|
Flag to enable or disable automatic saving. |
modified_on |
Optional[float]
|
Timestamp of the last modification. |
show_detailed_progress |
bool
|
Flag to show detailed progress during execution. |
show_edge_labels |
bool
|
Flag to show or hide named edge labels on connections. |
is_running |
bool
|
Indicates if the flow is currently running. |
is_canceled |
bool
|
Indicates if the flow execution has been canceled. |
track_history |
bool
|
Flag to enable or disable undo/redo history tracking. |
validate_settings |
bool
|
Flag to warn on nodes whose settings reference missing input columns. |
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
}
},
"description": "Extends FlowGraphConfig with additional operational settings for a flow.\n\nAttributes:\n auto_save (bool): Flag to enable or disable automatic saving.\n modified_on (Optional[float]): Timestamp of the last modification.\n show_detailed_progress (bool): Flag to show detailed progress during execution.\n show_edge_labels (bool): Flag to show or hide named edge labels on connections.\n is_running (bool): Indicates if the flow is currently running.\n is_canceled (bool): Indicates if the flow execution has been canceled.\n track_history (bool): Flag to enable or disable undo/redo history tracking.\n validate_settings (bool): Flag to warn on nodes whose settings reference missing input columns.",
"properties": {
"flow_id": {
"description": "Unique identifier for the flow.",
"title": "Flow Id",
"type": "integer"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"save_location": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Save Location"
},
"name": {
"default": "",
"title": "Name",
"type": "string"
},
"path": {
"default": "",
"title": "Path",
"type": "string"
},
"source_registration_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Catalog registration ID when running a registered flow.",
"title": "Source Registration Id"
},
"execution_mode": {
"default": "Performance",
"enum": [
"Development",
"Performance"
],
"title": "Execution Mode",
"type": "string"
},
"execution_location": {
"enum": [
"local",
"remote"
],
"title": "Execution Location",
"type": "string"
},
"max_parallel_workers": {
"default": 4,
"description": "Max threads for parallel node execution.",
"minimum": 1,
"title": "Max Parallel Workers",
"type": "integer"
},
"parameters": {
"description": "Flow-level parameters.",
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameters",
"type": "array"
},
"auto_save": {
"default": false,
"title": "Auto Save",
"type": "boolean"
},
"modified_on": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Modified On"
},
"show_detailed_progress": {
"default": true,
"title": "Show Detailed Progress",
"type": "boolean"
},
"show_edge_labels": {
"default": false,
"title": "Show Edge Labels",
"type": "boolean"
},
"is_running": {
"default": false,
"title": "Is Running",
"type": "boolean"
},
"is_canceled": {
"default": false,
"title": "Is Canceled",
"type": "boolean"
},
"track_history": {
"default": true,
"title": "Track History",
"type": "boolean"
},
"validate_settings": {
"default": true,
"title": "Validate Settings",
"type": "boolean"
}
},
"title": "FlowSettings",
"type": "object"
}
Fields:
-
flow_id(int) -
description(str | None) -
save_location(str | None) -
name(str) -
path(str) -
source_registration_id(int | None) -
execution_mode(ExecutionModeLiteral) -
execution_location(ExecutionLocationsLiteral) -
max_parallel_workers(int) -
parameters(list[FlowParameter]) -
auto_save(bool) -
modified_on(float | None) -
show_detailed_progress(bool) -
show_edge_labels(bool) -
is_running(bool) -
is_canceled(bool) -
track_history(bool) -
validate_settings(bool)
Validators:
-
validate_execution_mode→execution_mode -
validate_and_set_execution_location→execution_location
Source code in flowfile_core/flowfile_core/schemas/schemas.py
194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 | |
flow_id
pydantic-field
Unique identifier for the flow.
max_parallel_workers = 4
pydantic-field
Max threads for parallel node execution.
parameters
pydantic-field
Flow-level parameters.
source_registration_id = None
pydantic-field
Catalog registration ID when running a registered flow.
from_flow_settings_input(flow_graph_config)
classmethod
Creates a FlowSettings instance from a FlowGraphConfig instance.
:param flow_graph_config: The base flow graph configuration. :return: A new instance of FlowSettings with data from flow_graph_config.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
218 219 220 221 222 223 224 225 226 | |
validate_and_set_execution_location(v)
pydantic-validator
Validates and sets the execution location.
1. If None is provided: It defaults to the location determined by global settings.
2. If a value is provided: It checks if the value is compatible with the global
settings. If not (e.g., requesting 'remote' when only 'local' is possible),
it corrects the value to a compatible one.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | |
FlowSettingsResponse
pydantic-model
Bases: FlowSettings
FlowSettings plus runtime-only fields for API responses. Not persisted.
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
}
},
"description": "FlowSettings plus runtime-only fields for API responses. Not persisted.",
"properties": {
"flow_id": {
"description": "Unique identifier for the flow.",
"title": "Flow Id",
"type": "integer"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"save_location": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Save Location"
},
"name": {
"default": "",
"title": "Name",
"type": "string"
},
"path": {
"default": "",
"title": "Path",
"type": "string"
},
"source_registration_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Catalog registration ID when running a registered flow.",
"title": "Source Registration Id"
},
"execution_mode": {
"default": "Performance",
"enum": [
"Development",
"Performance"
],
"title": "Execution Mode",
"type": "string"
},
"execution_location": {
"enum": [
"local",
"remote"
],
"title": "Execution Location",
"type": "string"
},
"max_parallel_workers": {
"default": 4,
"description": "Max threads for parallel node execution.",
"minimum": 1,
"title": "Max Parallel Workers",
"type": "integer"
},
"parameters": {
"description": "Flow-level parameters.",
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameters",
"type": "array"
},
"auto_save": {
"default": false,
"title": "Auto Save",
"type": "boolean"
},
"modified_on": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Modified On"
},
"show_detailed_progress": {
"default": true,
"title": "Show Detailed Progress",
"type": "boolean"
},
"show_edge_labels": {
"default": false,
"title": "Show Edge Labels",
"type": "boolean"
},
"is_running": {
"default": false,
"title": "Is Running",
"type": "boolean"
},
"is_canceled": {
"default": false,
"title": "Is Canceled",
"type": "boolean"
},
"track_history": {
"default": true,
"title": "Track History",
"type": "boolean"
},
"validate_settings": {
"default": true,
"title": "Validate Settings",
"type": "boolean"
},
"has_unsaved_changes": {
"default": false,
"title": "Has Unsaved Changes",
"type": "boolean"
},
"display_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Display Name"
}
},
"title": "FlowSettingsResponse",
"type": "object"
}
Fields:
-
flow_id(int) -
description(str | None) -
save_location(str | None) -
name(str) -
path(str) -
source_registration_id(int | None) -
execution_mode(ExecutionModeLiteral) -
execution_location(ExecutionLocationsLiteral) -
max_parallel_workers(int) -
parameters(list[FlowParameter]) -
auto_save(bool) -
modified_on(float | None) -
show_detailed_progress(bool) -
show_edge_labels(bool) -
is_running(bool) -
is_canceled(bool) -
track_history(bool) -
validate_settings(bool) -
has_unsaved_changes(bool) -
display_name(str | None)
Validators:
-
validate_execution_mode→execution_mode -
validate_and_set_execution_location→execution_location
Source code in flowfile_core/flowfile_core/schemas/schemas.py
229 230 231 232 233 | |
flow_id
pydantic-field
Unique identifier for the flow.
max_parallel_workers = 4
pydantic-field
Max threads for parallel node execution.
parameters
pydantic-field
Flow-level parameters.
source_registration_id = None
pydantic-field
Catalog registration ID when running a registered flow.
from_flow_settings_input(flow_graph_config)
classmethod
Creates a FlowSettings instance from a FlowGraphConfig instance.
:param flow_graph_config: The base flow graph configuration. :return: A new instance of FlowSettings with data from flow_graph_config.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
218 219 220 221 222 223 224 225 226 | |
validate_and_set_execution_location(v)
pydantic-validator
Validates and sets the execution location.
1. If None is provided: It defaults to the location determined by global settings.
2. If a value is provided: It checks if the value is compatible with the global
settings. If not (e.g., requesting 'remote' when only 'local' is possible),
it corrects the value to a compatible one.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | |
FlowfileData
pydantic-model
Bases: BaseModel
Root model for flowfile serialization (YAML/JSON).
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
},
"FlowfileGroup": {
"description": "Serialized representation of a visual node group (YAML/JSON).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"default": 0.0,
"title": "X Position",
"type": "number"
},
"y_position": {
"default": 0.0,
"title": "Y Position",
"type": "number"
},
"width": {
"default": 400.0,
"title": "Width",
"type": "number"
},
"height": {
"default": 250.0,
"title": "Height",
"type": "number"
},
"collapsed": {
"default": false,
"title": "Collapsed",
"type": "boolean"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
}
},
"required": [
"id"
],
"title": "FlowfileGroup",
"type": "object"
},
"FlowfileInputConnection": {
"description": "One keyed input edge of a dynamic-input node (per-edge target handle).\n\nOnly nodes whose template sets ``dynamic_inputs`` serialize these; the same\nupstream node may legitimately appear twice with different handles.",
"properties": {
"from_id": {
"title": "From Id",
"type": "integer"
},
"input_handle": {
"title": "Input Handle",
"type": "string"
},
"source_handle": {
"default": "output-0",
"title": "Source Handle",
"type": "string"
}
},
"required": [
"from_id",
"input_handle"
],
"title": "FlowfileInputConnection",
"type": "object"
},
"FlowfileNode": {
"description": "Node representation for flowfile serialization (YAML/JSON).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"type": {
"title": "Type",
"type": "string"
},
"is_start_node": {
"default": false,
"title": "Is Start Node",
"type": "boolean"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"x_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "X Position"
},
"y_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "Y Position"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"left_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Input Id"
},
"right_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Input Id"
},
"input_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Input Ids"
},
"outputs": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Outputs"
},
"output_handles": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Handles"
},
"input_connections": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/FlowfileInputConnection"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Connections"
},
"setting_input": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Setting Input"
}
},
"required": [
"id",
"type"
],
"title": "FlowfileNode",
"type": "object"
},
"FlowfileSettings": {
"description": "Settings for flowfile serialization (YAML/JSON).\n\nExcludes runtime state fields like is_running, is_canceled, modified_on.",
"properties": {
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"execution_mode": {
"default": "Performance",
"enum": [
"Development",
"Performance"
],
"title": "Execution Mode",
"type": "string"
},
"execution_location": {
"default": "local",
"enum": [
"local",
"remote"
],
"title": "Execution Location",
"type": "string"
},
"auto_save": {
"default": false,
"title": "Auto Save",
"type": "boolean"
},
"show_detailed_progress": {
"default": true,
"title": "Show Detailed Progress",
"type": "boolean"
},
"validate_settings": {
"default": true,
"title": "Validate Settings",
"type": "boolean"
},
"max_parallel_workers": {
"default": 4,
"minimum": 1,
"title": "Max Parallel Workers",
"type": "integer"
},
"source_registration_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Registration Id"
},
"parameters": {
"description": "Flow-level parameters.",
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameters",
"type": "array"
}
},
"title": "FlowfileSettings",
"type": "object"
}
},
"description": "Root model for flowfile serialization (YAML/JSON).",
"properties": {
"flowfile_version": {
"title": "Flowfile Version",
"type": "string"
},
"flowfile_id": {
"title": "Flowfile Id",
"type": "integer"
},
"flowfile_name": {
"title": "Flowfile Name",
"type": "string"
},
"flowfile_settings": {
"$ref": "#/$defs/FlowfileSettings"
},
"nodes": {
"items": {
"$ref": "#/$defs/FlowfileNode"
},
"title": "Nodes",
"type": "array"
},
"groups": {
"items": {
"$ref": "#/$defs/FlowfileGroup"
},
"title": "Groups",
"type": "array"
}
},
"required": [
"flowfile_version",
"flowfile_id",
"flowfile_name",
"flowfile_settings",
"nodes"
],
"title": "FlowfileData",
"type": "object"
}
Fields:
-
flowfile_version(str) -
flowfile_id(int) -
flowfile_name(str) -
flowfile_settings(FlowfileSettings) -
nodes(list[FlowfileNode]) -
groups(list[FlowfileGroup])
Source code in flowfile_core/flowfile_core/schemas/schemas.py
387 388 389 390 391 392 393 394 395 | |
FlowfileGroup
pydantic-model
Bases: _GroupFields
Serialized representation of a visual node group (YAML/JSON).
Show JSON schema:
{
"description": "Serialized representation of a visual node group (YAML/JSON).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"default": 0.0,
"title": "X Position",
"type": "number"
},
"y_position": {
"default": 0.0,
"title": "Y Position",
"type": "number"
},
"width": {
"default": 400.0,
"title": "Width",
"type": "number"
},
"height": {
"default": 250.0,
"title": "Height",
"type": "number"
},
"collapsed": {
"default": false,
"title": "Collapsed",
"type": "boolean"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
}
},
"required": [
"id"
],
"title": "FlowfileGroup",
"type": "object"
}
Fields:
-
id(int) -
name(str) -
color(GroupColor | None) -
x_position(float) -
y_position(float) -
width(float) -
height(float) -
collapsed(bool) -
parent_group_id(int | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
383 384 | |
FlowfileInputConnection
pydantic-model
Bases: BaseModel
One keyed input edge of a dynamic-input node (per-edge target handle).
Only nodes whose template sets dynamic_inputs serialize these; the same
upstream node may legitimately appear twice with different handles.
Show JSON schema:
{
"description": "One keyed input edge of a dynamic-input node (per-edge target handle).\n\nOnly nodes whose template sets ``dynamic_inputs`` serialize these; the same\nupstream node may legitimately appear twice with different handles.",
"properties": {
"from_id": {
"title": "From Id",
"type": "integer"
},
"input_handle": {
"title": "Input Handle",
"type": "string"
},
"source_handle": {
"default": "output-0",
"title": "Source Handle",
"type": "string"
}
},
"required": [
"from_id",
"input_handle"
],
"title": "FlowfileInputConnection",
"type": "object"
}
Fields:
-
from_id(int) -
input_handle(str) -
source_handle(str)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
279 280 281 282 283 284 285 286 287 288 | |
FlowfileNode
pydantic-model
Bases: BaseModel
Node representation for flowfile serialization (YAML/JSON).
Show JSON schema:
{
"$defs": {
"FlowfileInputConnection": {
"description": "One keyed input edge of a dynamic-input node (per-edge target handle).\n\nOnly nodes whose template sets ``dynamic_inputs`` serialize these; the same\nupstream node may legitimately appear twice with different handles.",
"properties": {
"from_id": {
"title": "From Id",
"type": "integer"
},
"input_handle": {
"title": "Input Handle",
"type": "string"
},
"source_handle": {
"default": "output-0",
"title": "Source Handle",
"type": "string"
}
},
"required": [
"from_id",
"input_handle"
],
"title": "FlowfileInputConnection",
"type": "object"
}
},
"description": "Node representation for flowfile serialization (YAML/JSON).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"type": {
"title": "Type",
"type": "string"
},
"is_start_node": {
"default": false,
"title": "Is Start Node",
"type": "boolean"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"x_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "X Position"
},
"y_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "Y Position"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"left_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Input Id"
},
"right_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Input Id"
},
"input_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Input Ids"
},
"outputs": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Outputs"
},
"output_handles": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Handles"
},
"input_connections": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/FlowfileInputConnection"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Connections"
},
"setting_input": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Setting Input"
}
},
"required": [
"id",
"type"
],
"title": "FlowfileNode",
"type": "object"
}
Fields:
-
id(int) -
type(str) -
is_start_node(bool) -
description(str | None) -
node_reference(str | None) -
x_position(int | None) -
y_position(int | None) -
group_id(int | None) -
left_input_id(int | None) -
right_input_id(int | None) -
input_ids(list[int] | None) -
outputs(list[int] | None) -
output_handles(list[str] | None) -
input_connections(list[FlowfileInputConnection] | None) -
setting_input(Any | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 | |
FlowfileSettings
pydantic-model
Bases: BaseModel
Settings for flowfile serialization (YAML/JSON).
Excludes runtime state fields like is_running, is_canceled, modified_on.
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
}
},
"description": "Settings for flowfile serialization (YAML/JSON).\n\nExcludes runtime state fields like is_running, is_canceled, modified_on.",
"properties": {
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"execution_mode": {
"default": "Performance",
"enum": [
"Development",
"Performance"
],
"title": "Execution Mode",
"type": "string"
},
"execution_location": {
"default": "local",
"enum": [
"local",
"remote"
],
"title": "Execution Location",
"type": "string"
},
"auto_save": {
"default": false,
"title": "Auto Save",
"type": "boolean"
},
"show_detailed_progress": {
"default": true,
"title": "Show Detailed Progress",
"type": "boolean"
},
"validate_settings": {
"default": true,
"title": "Validate Settings",
"type": "boolean"
},
"max_parallel_workers": {
"default": 4,
"minimum": 1,
"title": "Max Parallel Workers",
"type": "integer"
},
"source_registration_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Registration Id"
},
"parameters": {
"description": "Flow-level parameters.",
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameters",
"type": "array"
}
},
"title": "FlowfileSettings",
"type": "object"
}
Fields:
-
description(str | None) -
execution_mode(ExecutionModeLiteral) -
execution_location(ExecutionLocationsLiteral) -
auto_save(bool) -
show_detailed_progress(bool) -
validate_settings(bool) -
max_parallel_workers(int) -
source_registration_id(int | None) -
parameters(list[FlowParameter])
Validators:
-
validate_execution_mode→execution_mode
Source code in flowfile_core/flowfile_core/schemas/schemas.py
255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 | |
parameters
pydantic-field
Flow-level parameters.
GroupBounds
Bases: NamedTuple
Axis-aligned bounds of a group box, in absolute canvas coordinates.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
353 354 355 356 357 358 359 | |
GroupBoundsUpdate
pydantic-model
Bases: BaseModel
A single group's new absolute bounds.
Show JSON schema:
{
"description": "A single group's new absolute bounds.",
"properties": {
"group_id": {
"title": "Group Id",
"type": "integer"
},
"x_position": {
"title": "X Position",
"type": "number"
},
"y_position": {
"title": "Y Position",
"type": "number"
},
"width": {
"title": "Width",
"type": "number"
},
"height": {
"title": "Height",
"type": "number"
}
},
"required": [
"group_id",
"x_position",
"y_position",
"width",
"height"
],
"title": "GroupBoundsUpdate",
"type": "object"
}
Fields:
-
group_id(int) -
x_position(float) -
y_position(float) -
width(float) -
height(float)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
847 848 849 850 851 852 853 854 | |
GroupInformation
pydantic-model
Bases: _GroupFields
Runtime representation of a visual node group (stored in FlowGraph._groups).
Show JSON schema:
{
"description": "Runtime representation of a visual node group (stored in FlowGraph._groups).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"default": 0.0,
"title": "X Position",
"type": "number"
},
"y_position": {
"default": 0.0,
"title": "Y Position",
"type": "number"
},
"width": {
"default": 400.0,
"title": "Width",
"type": "number"
},
"height": {
"default": 250.0,
"title": "Height",
"type": "number"
},
"collapsed": {
"default": false,
"title": "Collapsed",
"type": "boolean"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
}
},
"required": [
"id"
],
"title": "GroupInformation",
"type": "object"
}
Fields:
-
id(int) -
name(str) -
color(GroupColor | None) -
x_position(float) -
y_position(float) -
width(float) -
height(float) -
collapsed(bool) -
parent_group_id(int | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
379 380 | |
GroupMembershipRequest
pydantic-model
Bases: BaseModel
Body for adding/removing nodes from a group.
Show JSON schema:
{
"description": "Body for adding/removing nodes from a group.",
"properties": {
"node_ids": {
"items": {
"type": "integer"
},
"title": "Node Ids",
"type": "array"
}
},
"required": [
"node_ids"
],
"title": "GroupMembershipRequest",
"type": "object"
}
Fields:
-
node_ids(list[int])
Source code in flowfile_core/flowfile_core/schemas/schemas.py
833 834 835 836 | |
NodeConnection
pydantic-model
Bases: BaseModel
Represents a connection between two nodes in the flow.
Attributes:
| Name | Type | Description |
|---|---|---|
from_node_id |
int
|
The ID of the source node. |
to_node_id |
int
|
The ID of the target node. |
Show JSON schema:
{
"description": "Represents a connection between two nodes in the flow.\n\nAttributes:\n from_node_id (int): The ID of the source node.\n to_node_id (int): The ID of the target node.",
"properties": {
"from_node_id": {
"title": "From Node Id",
"type": "integer"
},
"to_node_id": {
"title": "To Node Id",
"type": "integer"
}
},
"required": [
"from_node_id",
"to_node_id"
],
"title": "NodeConnection",
"type": "object"
}
Config:
frozen:True
Fields:
-
from_node_id(int) -
to_node_id(int)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
730 731 732 733 734 735 736 737 738 739 740 741 | |
NodeDefault
pydantic-model
Bases: BaseModel
Defines default properties for a node type.
Attributes:
| Name | Type | Description |
|---|---|---|
node_name |
str
|
The name of the node. |
node_type |
NodeTypeLiteral
|
The functional type of the node ('input', 'output', 'process'). |
transform_type |
TransformTypeLiteral
|
The data transformation behavior ('narrow', 'wide', 'other'). |
has_default_settings |
Optional[Any]
|
Indicates if the node has predefined default settings. |
Show JSON schema:
{
"description": "Defines default properties for a node type.\n\nAttributes:\n node_name (str): The name of the node.\n node_type (NodeTypeLiteral): The functional type of the node ('input', 'output', 'process').\n transform_type (TransformTypeLiteral): The data transformation behavior ('narrow', 'wide', 'other').\n has_default_settings (Optional[Any]): Indicates if the node has predefined default settings.",
"properties": {
"node_name": {
"title": "Node Name",
"type": "string"
},
"node_type": {
"enum": [
"input",
"output",
"process"
],
"title": "Node Type",
"type": "string"
},
"transform_type": {
"enum": [
"narrow",
"wide",
"other"
],
"title": "Transform Type",
"type": "string"
},
"has_default_settings": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Has Default Settings"
}
},
"required": [
"node_name",
"node_type",
"transform_type"
],
"title": "NodeDefault",
"type": "object"
}
Fields:
-
node_name(str) -
node_type(NodeTypeLiteral) -
transform_type(TransformTypeLiteral) -
has_default_settings(Any | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 | |
NodeEdge
pydantic-model
Bases: BaseModel
Represents a connection (edge) between two nodes in the frontend.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
A unique identifier for the edge. |
source |
str
|
The ID of the source node. |
target |
str
|
The ID of the target node. |
targetHandle |
str
|
The specific input handle on the target node. |
sourceHandle |
str
|
The specific output handle on the source node. |
Show JSON schema:
{
"description": "Represents a connection (edge) between two nodes in the frontend.\n\nAttributes:\n id (str): A unique identifier for the edge.\n source (str): The ID of the source node.\n target (str): The ID of the target node.\n targetHandle (str): The specific input handle on the target node.\n sourceHandle (str): The specific output handle on the source node.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"source": {
"title": "Source",
"type": "string"
},
"target": {
"title": "Target",
"type": "string"
},
"targetHandle": {
"title": "Targethandle",
"type": "string"
},
"sourceHandle": {
"title": "Sourcehandle",
"type": "string"
}
},
"required": [
"id",
"source",
"target",
"targetHandle",
"sourceHandle"
],
"title": "NodeEdge",
"type": "object"
}
Config:
coerce_numbers_to_str:True
Fields:
-
id(str) -
source(str) -
target(str) -
targetHandle(str) -
sourceHandle(str)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 | |
NodeInformation
pydantic-model
Bases: BaseModel
Stores the state and configuration of a specific node instance within a flow.
Show JSON schema:
{
"$defs": {
"FlowfileInputConnection": {
"description": "One keyed input edge of a dynamic-input node (per-edge target handle).\n\nOnly nodes whose template sets ``dynamic_inputs`` serialize these; the same\nupstream node may legitimately appear twice with different handles.",
"properties": {
"from_id": {
"title": "From Id",
"type": "integer"
},
"input_handle": {
"title": "Input Handle",
"type": "string"
},
"source_handle": {
"default": "output-0",
"title": "Source Handle",
"type": "string"
}
},
"required": [
"from_id",
"input_handle"
],
"title": "FlowfileInputConnection",
"type": "object"
}
},
"description": "Stores the state and configuration of a specific node instance within a flow.",
"properties": {
"id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Id"
},
"type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Type"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Is Setup"
},
"is_start_node": {
"default": false,
"title": "Is Start Node",
"type": "boolean"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"x_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "X Position"
},
"y_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": 0,
"title": "Y Position"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"left_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Input Id"
},
"right_input_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Input Id"
},
"input_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Input Ids"
},
"outputs": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Outputs"
},
"output_handles": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Handles"
},
"input_connections": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/FlowfileInputConnection"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Connections"
},
"setting_input": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Setting Input"
}
},
"title": "NodeInformation",
"type": "object"
}
Fields:
-
id(int | None) -
type(str | None) -
is_setup(bool | None) -
is_start_node(bool) -
description(str | None) -
node_reference(str | None) -
x_position(int | None) -
y_position(int | None) -
group_id(int | None) -
left_input_id(int | None) -
right_input_id(int | None) -
input_ids(list[int] | None) -
outputs(list[int] | None) -
output_handles(list[str] | None) -
input_connections(list[FlowfileInputConnection] | None) -
setting_input(Any | None)
Validators:
-
validate_setting_input→setting_input
Source code in flowfile_core/flowfile_core/schemas/schemas.py
649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 | |
NodeInput
pydantic-model
Bases: NodeTemplate
Represents a node as it is received from the frontend, including position.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
int
|
The unique ID of the node instance. |
pos_x |
float
|
The x-coordinate on the canvas. |
pos_y |
float
|
The y-coordinate on the canvas. |
output_names |
list[str] | None
|
Named outputs for multi-output nodes. |
node_reference |
str | None
|
Reference name used for code generation and input naming. |
Show JSON schema:
{
"$defs": {
"ArtifactDecl": {
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Type"
}
},
"required": [
"name"
],
"title": "ArtifactDecl",
"type": "object"
},
"NodeTag": {
"description": "Controlled vocabulary of palette search keywords.\n\nMatched (case-insensitive substring) against the user's query in the node palette so a\nnode surfaces by concept, format, or tool rather than only its display name\n(e.g. \"s3\" -> cloud reader/writer, \"sum\" -> formula and group by). As a ``str`` enum each\nmember serializes to its plain string value for the frontend.",
"enum": [
"csv",
"excel",
"parquet",
"json",
"file",
"read",
"write",
"import",
"export",
"save",
"delta",
"api",
"rest",
"http",
"external",
"response",
"pagination",
"database",
"sql",
"query",
"table",
"postgres",
"mysql",
"sql server",
"snowflake",
"oracle",
"sqlite",
"redshift",
"bigquery",
"s3",
"aws",
"azure",
"adls",
"gcs",
"blob",
"bucket",
"cloud",
"catalog",
"lakehouse",
"time travel",
"kafka",
"redpanda",
"streaming",
"topic",
"google analytics",
"ga4",
"analytics",
"manual",
"paste",
"input",
"select",
"columns",
"rename",
"reorder",
"projection",
"filter",
"where",
"subset",
"sample",
"limit",
"head",
"formula",
"expression",
"calculate",
"math",
"concat",
"transform",
"group by",
"aggregate",
"sum",
"mean",
"average",
"count",
"min",
"max",
"median",
"summarize",
"record count",
"rows",
"window",
"rolling",
"cumulative",
"rank",
"partition",
"lag",
"lead",
"join",
"merge",
"lookup",
"vlookup",
"inner",
"outer",
"cross join",
"cartesian",
"fuzzy",
"similarity",
"levenshtein",
"union",
"append",
"wait",
"dependency",
"pivot",
"crosstab",
"unpivot",
"melt",
"reshape",
"text to rows",
"split",
"explode",
"unique",
"dedupe",
"distinct",
"drop duplicates",
"graph",
"network",
"cluster",
"connected components",
"record id",
"row number",
"index",
"sort",
"order",
"ascending",
"descending",
"polars",
"code",
"python",
"script",
"kernel",
"custom",
"dataframe",
"explore",
"profile",
"preview",
"eda",
"statistics",
"visualize",
"bar chart",
"insight",
"graphs",
"ml",
"machine learning",
"train",
"test",
"model",
"regression",
"classification",
"predict",
"score",
"evaluate",
"metrics"
],
"title": "NodeTag",
"type": "string"
}
},
"description": "Represents a node as it is received from the frontend, including position.\n\nAttributes:\n id (int): The unique ID of the node instance.\n pos_x (float): The x-coordinate on the canvas.\n pos_y (float): The y-coordinate on the canvas.\n output_names (list[str] | None): Named outputs for multi-output nodes.\n node_reference (str | None): Reference name used for code generation and input naming.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"item": {
"title": "Item",
"type": "string"
},
"input": {
"title": "Input",
"type": "integer"
},
"output": {
"title": "Output",
"type": "integer"
},
"image": {
"title": "Image",
"type": "string"
},
"multi": {
"default": false,
"title": "Multi",
"type": "boolean"
},
"node_type": {
"enum": [
"input",
"output",
"process"
],
"title": "Node Type",
"type": "string"
},
"transform_type": {
"enum": [
"narrow",
"wide",
"other"
],
"title": "Transform Type",
"type": "string"
},
"node_group": {
"title": "Node Group",
"type": "string"
},
"node_group_label": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Group Label"
},
"prod_ready": {
"default": true,
"title": "Prod Ready",
"type": "boolean"
},
"can_be_start": {
"default": false,
"title": "Can Be Start",
"type": "boolean"
},
"drawer_title": {
"default": "Node title",
"title": "Drawer Title",
"type": "string"
},
"drawer_intro": {
"default": "Drawer into",
"title": "Drawer Intro",
"type": "string"
},
"custom_node": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Custom Node"
},
"execution_environment": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Execution Environment"
},
"dependencies": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Dependencies"
},
"publishes": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/ArtifactDecl"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Publishes"
},
"laziness": {
"default": "eager",
"enum": [
"lazy",
"eager",
"conditional"
],
"title": "Laziness",
"type": "string"
},
"output_names": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Names"
},
"input_labels": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Labels"
},
"dynamic_inputs": {
"default": false,
"title": "Dynamic Inputs",
"type": "boolean"
},
"tags": {
"items": {
"$ref": "#/$defs/NodeTag"
},
"title": "Tags",
"type": "array"
},
"id": {
"title": "Id",
"type": "integer"
},
"pos_x": {
"title": "Pos X",
"type": "number"
},
"pos_y": {
"title": "Pos Y",
"type": "number"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"input_names": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Names"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
}
},
"required": [
"name",
"item",
"input",
"output",
"image",
"node_type",
"transform_type",
"node_group",
"id",
"pos_x",
"pos_y"
],
"title": "NodeInput",
"type": "object"
}
Fields:
-
name(str) -
item(str) -
input(int) -
output(int) -
image(str) -
multi(bool) -
node_type(NodeTypeLiteral) -
transform_type(TransformTypeLiteral) -
node_group(str) -
node_group_label(str | None) -
prod_ready(bool) -
can_be_start(bool) -
drawer_title(str) -
drawer_intro(str) -
custom_node(bool | None) -
execution_environment(str | None) -
dependencies(list[str] | None) -
publishes(list[ArtifactDecl] | None) -
laziness(LazinessLiteral) -
input_labels(list[str] | None) -
dynamic_inputs(bool) -
tags(list[NodeTag]) -
id(int) -
pos_x(float) -
pos_y(float) -
group_id(int | None) -
output_names(list[str] | None) -
input_names(list[str] | None) -
node_reference(str | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 | |
NodePositionUpdate
pydantic-model
Bases: BaseModel
A single node's new absolute canvas position.
Show JSON schema:
{
"description": "A single node's new absolute canvas position.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"pos_x": {
"title": "Pos X",
"type": "number"
},
"pos_y": {
"title": "Pos Y",
"type": "number"
}
},
"required": [
"node_id",
"pos_x",
"pos_y"
],
"title": "NodePositionUpdate",
"type": "object"
}
Fields:
-
node_id(int) -
pos_x(float) -
pos_y(float)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
839 840 841 842 843 844 | |
NodeTag
Bases: str, Enum
Controlled vocabulary of palette search keywords.
Matched (case-insensitive substring) against the user's query in the node palette so a
node surfaces by concept, format, or tool rather than only its display name
(e.g. "s3" -> cloud reader/writer, "sum" -> formula and group by). As a str enum each
member serializes to its plain string value for the frontend.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 | |
NodeTemplate
pydantic-model
Bases: BaseModel
Defines the template for a node type, specifying its UI and functional characteristics.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
The display name of the node. |
item |
str
|
The unique identifier for the node type. |
input |
int
|
The number of required input connections. |
output |
int
|
The number of output connections. |
image |
str
|
The filename of the icon for the node. |
multi |
bool
|
Whether the node accepts multiple main input connections. |
node_group |
str
|
The category group the node belongs to (e.g., 'input', 'transform'). |
prod_ready |
bool
|
Whether the node is considered production-ready. |
can_be_start |
bool
|
Whether the node can be a starting point in a flow. |
Show JSON schema:
{
"$defs": {
"ArtifactDecl": {
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Type"
}
},
"required": [
"name"
],
"title": "ArtifactDecl",
"type": "object"
},
"NodeTag": {
"description": "Controlled vocabulary of palette search keywords.\n\nMatched (case-insensitive substring) against the user's query in the node palette so a\nnode surfaces by concept, format, or tool rather than only its display name\n(e.g. \"s3\" -> cloud reader/writer, \"sum\" -> formula and group by). As a ``str`` enum each\nmember serializes to its plain string value for the frontend.",
"enum": [
"csv",
"excel",
"parquet",
"json",
"file",
"read",
"write",
"import",
"export",
"save",
"delta",
"api",
"rest",
"http",
"external",
"response",
"pagination",
"database",
"sql",
"query",
"table",
"postgres",
"mysql",
"sql server",
"snowflake",
"oracle",
"sqlite",
"redshift",
"bigquery",
"s3",
"aws",
"azure",
"adls",
"gcs",
"blob",
"bucket",
"cloud",
"catalog",
"lakehouse",
"time travel",
"kafka",
"redpanda",
"streaming",
"topic",
"google analytics",
"ga4",
"analytics",
"manual",
"paste",
"input",
"select",
"columns",
"rename",
"reorder",
"projection",
"filter",
"where",
"subset",
"sample",
"limit",
"head",
"formula",
"expression",
"calculate",
"math",
"concat",
"transform",
"group by",
"aggregate",
"sum",
"mean",
"average",
"count",
"min",
"max",
"median",
"summarize",
"record count",
"rows",
"window",
"rolling",
"cumulative",
"rank",
"partition",
"lag",
"lead",
"join",
"merge",
"lookup",
"vlookup",
"inner",
"outer",
"cross join",
"cartesian",
"fuzzy",
"similarity",
"levenshtein",
"union",
"append",
"wait",
"dependency",
"pivot",
"crosstab",
"unpivot",
"melt",
"reshape",
"text to rows",
"split",
"explode",
"unique",
"dedupe",
"distinct",
"drop duplicates",
"graph",
"network",
"cluster",
"connected components",
"record id",
"row number",
"index",
"sort",
"order",
"ascending",
"descending",
"polars",
"code",
"python",
"script",
"kernel",
"custom",
"dataframe",
"explore",
"profile",
"preview",
"eda",
"statistics",
"visualize",
"bar chart",
"insight",
"graphs",
"ml",
"machine learning",
"train",
"test",
"model",
"regression",
"classification",
"predict",
"score",
"evaluate",
"metrics"
],
"title": "NodeTag",
"type": "string"
}
},
"description": "Defines the template for a node type, specifying its UI and functional characteristics.\n\nAttributes:\n name (str): The display name of the node.\n item (str): The unique identifier for the node type.\n input (int): The number of required input connections.\n output (int): The number of output connections.\n image (str): The filename of the icon for the node.\n multi (bool): Whether the node accepts multiple main input connections.\n node_group (str): The category group the node belongs to (e.g., 'input', 'transform').\n prod_ready (bool): Whether the node is considered production-ready.\n can_be_start (bool): Whether the node can be a starting point in a flow.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"item": {
"title": "Item",
"type": "string"
},
"input": {
"title": "Input",
"type": "integer"
},
"output": {
"title": "Output",
"type": "integer"
},
"image": {
"title": "Image",
"type": "string"
},
"multi": {
"default": false,
"title": "Multi",
"type": "boolean"
},
"node_type": {
"enum": [
"input",
"output",
"process"
],
"title": "Node Type",
"type": "string"
},
"transform_type": {
"enum": [
"narrow",
"wide",
"other"
],
"title": "Transform Type",
"type": "string"
},
"node_group": {
"title": "Node Group",
"type": "string"
},
"node_group_label": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Group Label"
},
"prod_ready": {
"default": true,
"title": "Prod Ready",
"type": "boolean"
},
"can_be_start": {
"default": false,
"title": "Can Be Start",
"type": "boolean"
},
"drawer_title": {
"default": "Node title",
"title": "Drawer Title",
"type": "string"
},
"drawer_intro": {
"default": "Drawer into",
"title": "Drawer Intro",
"type": "string"
},
"custom_node": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Custom Node"
},
"execution_environment": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Execution Environment"
},
"dependencies": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Dependencies"
},
"publishes": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/ArtifactDecl"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Publishes"
},
"laziness": {
"default": "eager",
"enum": [
"lazy",
"eager",
"conditional"
],
"title": "Laziness",
"type": "string"
},
"output_names": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Names"
},
"input_labels": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Labels"
},
"dynamic_inputs": {
"default": false,
"title": "Dynamic Inputs",
"type": "boolean"
},
"tags": {
"items": {
"$ref": "#/$defs/NodeTag"
},
"title": "Tags",
"type": "array"
}
},
"required": [
"name",
"item",
"input",
"output",
"image",
"node_type",
"transform_type",
"node_group"
],
"title": "NodeTemplate",
"type": "object"
}
Fields:
-
name(str) -
item(str) -
input(int) -
output(int) -
image(str) -
multi(bool) -
node_type(NodeTypeLiteral) -
transform_type(TransformTypeLiteral) -
node_group(str) -
node_group_label(str | None) -
prod_ready(bool) -
can_be_start(bool) -
drawer_title(str) -
drawer_intro(str) -
custom_node(bool | None) -
execution_environment(str | None) -
dependencies(list[str] | None) -
publishes(list[ArtifactDecl] | None) -
laziness(LazinessLiteral) -
output_names(list[str] | None) -
input_labels(list[str] | None) -
dynamic_inputs(bool) -
tags(list[NodeTag])
Source code in flowfile_core/flowfile_core/schemas/schemas.py
601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 | |
RawLogInput
pydantic-model
Bases: BaseModel
Schema for a raw log message.
Attributes:
| Name | Type | Description |
|---|---|---|
flowfile_flow_id |
int
|
The ID of the flow that generated the log. |
log_message |
str
|
The content of the log message. |
log_type |
Literal['INFO', 'WARNING', 'ERROR']
|
The type of log. |
node_id |
int | None
|
Optional node ID to attribute the log to. |
extra |
Optional[dict]
|
Extra context data for the log. |
Show JSON schema:
{
"description": "Schema for a raw log message.\n\nAttributes:\n flowfile_flow_id (int): The ID of the flow that generated the log.\n log_message (str): The content of the log message.\n log_type (Literal[\"INFO\", \"WARNING\", \"ERROR\"]): The type of log.\n node_id (int | None): Optional node ID to attribute the log to.\n extra (Optional[dict]): Extra context data for the log.",
"properties": {
"flowfile_flow_id": {
"title": "Flowfile Flow Id",
"type": "integer"
},
"log_message": {
"title": "Log Message",
"type": "string"
},
"log_type": {
"enum": [
"INFO",
"WARNING",
"ERROR"
],
"title": "Log Type",
"type": "string"
},
"node_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Id"
},
"extra": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Extra"
}
},
"required": [
"flowfile_flow_id",
"log_message",
"log_type"
],
"title": "RawLogInput",
"type": "object"
}
Fields:
-
flowfile_flow_id(int) -
log_message(str) -
log_type(Literal['INFO', 'WARNING', 'ERROR']) -
node_id(int | None) -
extra(dict | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 | |
UpdateGroupRequest
pydantic-model
Bases: BaseModel
Body for POST /editor/update_group/. All fields optional -> partial update.
Show JSON schema:
{
"description": "Body for POST /editor/update_group/. All fields optional -> partial update.",
"properties": {
"name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Name"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "X Position"
},
"y_position": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Y Position"
},
"width": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Width"
},
"height": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"title": "Height"
},
"collapsed": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Collapsed"
}
},
"title": "UpdateGroupRequest",
"type": "object"
}
Fields:
-
name(str | None) -
color(GroupColor | None) -
x_position(float | None) -
y_position(float | None) -
width(float | None) -
height(float | None) -
collapsed(bool | None)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
821 822 823 824 825 826 827 828 829 830 | |
UpdateLayoutRequest
pydantic-model
Bases: BaseModel
Batch persistence of dragged node positions and/or group bounds (one drag-end -> one call).
Show JSON schema:
{
"$defs": {
"GroupBoundsUpdate": {
"description": "A single group's new absolute bounds.",
"properties": {
"group_id": {
"title": "Group Id",
"type": "integer"
},
"x_position": {
"title": "X Position",
"type": "number"
},
"y_position": {
"title": "Y Position",
"type": "number"
},
"width": {
"title": "Width",
"type": "number"
},
"height": {
"title": "Height",
"type": "number"
}
},
"required": [
"group_id",
"x_position",
"y_position",
"width",
"height"
],
"title": "GroupBoundsUpdate",
"type": "object"
},
"NodePositionUpdate": {
"description": "A single node's new absolute canvas position.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"pos_x": {
"title": "Pos X",
"type": "number"
},
"pos_y": {
"title": "Pos Y",
"type": "number"
}
},
"required": [
"node_id",
"pos_x",
"pos_y"
],
"title": "NodePositionUpdate",
"type": "object"
}
},
"description": "Batch persistence of dragged node positions and/or group bounds (one drag-end -> one call).",
"properties": {
"node_positions": {
"items": {
"$ref": "#/$defs/NodePositionUpdate"
},
"title": "Node Positions",
"type": "array"
},
"group_bounds": {
"items": {
"$ref": "#/$defs/GroupBoundsUpdate"
},
"title": "Group Bounds",
"type": "array"
},
"record_history": {
"default": true,
"title": "Record History",
"type": "boolean"
}
},
"title": "UpdateLayoutRequest",
"type": "object"
}
Fields:
-
node_positions(list[NodePositionUpdate]) -
group_bounds(list[GroupBoundsUpdate]) -
record_history(bool)
Source code in flowfile_core/flowfile_core/schemas/schemas.py
857 858 859 860 861 862 863 | |
VueFlowInput
pydantic-model
Bases: BaseModel
Represents the complete graph structure from the Vue-based frontend.
Attributes:
| Name | Type | Description |
|---|---|---|
node_edges |
List[NodeEdge]
|
A list of all edges in the graph. |
node_inputs |
List[NodeInput]
|
A list of all nodes in the graph. |
Show JSON schema:
{
"$defs": {
"ArtifactDecl": {
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Type"
}
},
"required": [
"name"
],
"title": "ArtifactDecl",
"type": "object"
},
"FlowfileGroup": {
"description": "Serialized representation of a visual node group (YAML/JSON).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"default": 0.0,
"title": "X Position",
"type": "number"
},
"y_position": {
"default": 0.0,
"title": "Y Position",
"type": "number"
},
"width": {
"default": 400.0,
"title": "Width",
"type": "number"
},
"height": {
"default": 250.0,
"title": "Height",
"type": "number"
},
"collapsed": {
"default": false,
"title": "Collapsed",
"type": "boolean"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
}
},
"required": [
"id"
],
"title": "FlowfileGroup",
"type": "object"
},
"NodeEdge": {
"description": "Represents a connection (edge) between two nodes in the frontend.\n\nAttributes:\n id (str): A unique identifier for the edge.\n source (str): The ID of the source node.\n target (str): The ID of the target node.\n targetHandle (str): The specific input handle on the target node.\n sourceHandle (str): The specific output handle on the source node.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"source": {
"title": "Source",
"type": "string"
},
"target": {
"title": "Target",
"type": "string"
},
"targetHandle": {
"title": "Targethandle",
"type": "string"
},
"sourceHandle": {
"title": "Sourcehandle",
"type": "string"
}
},
"required": [
"id",
"source",
"target",
"targetHandle",
"sourceHandle"
],
"title": "NodeEdge",
"type": "object"
},
"NodeInput": {
"description": "Represents a node as it is received from the frontend, including position.\n\nAttributes:\n id (int): The unique ID of the node instance.\n pos_x (float): The x-coordinate on the canvas.\n pos_y (float): The y-coordinate on the canvas.\n output_names (list[str] | None): Named outputs for multi-output nodes.\n node_reference (str | None): Reference name used for code generation and input naming.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"item": {
"title": "Item",
"type": "string"
},
"input": {
"title": "Input",
"type": "integer"
},
"output": {
"title": "Output",
"type": "integer"
},
"image": {
"title": "Image",
"type": "string"
},
"multi": {
"default": false,
"title": "Multi",
"type": "boolean"
},
"node_type": {
"enum": [
"input",
"output",
"process"
],
"title": "Node Type",
"type": "string"
},
"transform_type": {
"enum": [
"narrow",
"wide",
"other"
],
"title": "Transform Type",
"type": "string"
},
"node_group": {
"title": "Node Group",
"type": "string"
},
"node_group_label": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Group Label"
},
"prod_ready": {
"default": true,
"title": "Prod Ready",
"type": "boolean"
},
"can_be_start": {
"default": false,
"title": "Can Be Start",
"type": "boolean"
},
"drawer_title": {
"default": "Node title",
"title": "Drawer Title",
"type": "string"
},
"drawer_intro": {
"default": "Drawer into",
"title": "Drawer Intro",
"type": "string"
},
"custom_node": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Custom Node"
},
"execution_environment": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Execution Environment"
},
"dependencies": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Dependencies"
},
"publishes": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/ArtifactDecl"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Publishes"
},
"laziness": {
"default": "eager",
"enum": [
"lazy",
"eager",
"conditional"
],
"title": "Laziness",
"type": "string"
},
"output_names": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Names"
},
"input_labels": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Labels"
},
"dynamic_inputs": {
"default": false,
"title": "Dynamic Inputs",
"type": "boolean"
},
"tags": {
"items": {
"$ref": "#/$defs/NodeTag"
},
"title": "Tags",
"type": "array"
},
"id": {
"title": "Id",
"type": "integer"
},
"pos_x": {
"title": "Pos X",
"type": "number"
},
"pos_y": {
"title": "Pos Y",
"type": "number"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"input_names": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Input Names"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
}
},
"required": [
"name",
"item",
"input",
"output",
"image",
"node_type",
"transform_type",
"node_group",
"id",
"pos_x",
"pos_y"
],
"title": "NodeInput",
"type": "object"
},
"NodeTag": {
"description": "Controlled vocabulary of palette search keywords.\n\nMatched (case-insensitive substring) against the user's query in the node palette so a\nnode surfaces by concept, format, or tool rather than only its display name\n(e.g. \"s3\" -> cloud reader/writer, \"sum\" -> formula and group by). As a ``str`` enum each\nmember serializes to its plain string value for the frontend.",
"enum": [
"csv",
"excel",
"parquet",
"json",
"file",
"read",
"write",
"import",
"export",
"save",
"delta",
"api",
"rest",
"http",
"external",
"response",
"pagination",
"database",
"sql",
"query",
"table",
"postgres",
"mysql",
"sql server",
"snowflake",
"oracle",
"sqlite",
"redshift",
"bigquery",
"s3",
"aws",
"azure",
"adls",
"gcs",
"blob",
"bucket",
"cloud",
"catalog",
"lakehouse",
"time travel",
"kafka",
"redpanda",
"streaming",
"topic",
"google analytics",
"ga4",
"analytics",
"manual",
"paste",
"input",
"select",
"columns",
"rename",
"reorder",
"projection",
"filter",
"where",
"subset",
"sample",
"limit",
"head",
"formula",
"expression",
"calculate",
"math",
"concat",
"transform",
"group by",
"aggregate",
"sum",
"mean",
"average",
"count",
"min",
"max",
"median",
"summarize",
"record count",
"rows",
"window",
"rolling",
"cumulative",
"rank",
"partition",
"lag",
"lead",
"join",
"merge",
"lookup",
"vlookup",
"inner",
"outer",
"cross join",
"cartesian",
"fuzzy",
"similarity",
"levenshtein",
"union",
"append",
"wait",
"dependency",
"pivot",
"crosstab",
"unpivot",
"melt",
"reshape",
"text to rows",
"split",
"explode",
"unique",
"dedupe",
"distinct",
"drop duplicates",
"graph",
"network",
"cluster",
"connected components",
"record id",
"row number",
"index",
"sort",
"order",
"ascending",
"descending",
"polars",
"code",
"python",
"script",
"kernel",
"custom",
"dataframe",
"explore",
"profile",
"preview",
"eda",
"statistics",
"visualize",
"bar chart",
"insight",
"graphs",
"ml",
"machine learning",
"train",
"test",
"model",
"regression",
"classification",
"predict",
"score",
"evaluate",
"metrics"
],
"title": "NodeTag",
"type": "string"
}
},
"description": "Represents the complete graph structure from the Vue-based frontend.\n\nAttributes:\n node_edges (List[NodeEdge]): A list of all edges in the graph.\n node_inputs (List[NodeInput]): A list of all nodes in the graph.",
"properties": {
"node_edges": {
"items": {
"$ref": "#/$defs/NodeEdge"
},
"title": "Node Edges",
"type": "array"
},
"node_inputs": {
"items": {
"$ref": "#/$defs/NodeInput"
},
"title": "Node Inputs",
"type": "array"
},
"groups": {
"items": {
"$ref": "#/$defs/FlowfileGroup"
},
"title": "Groups",
"type": "array"
}
},
"required": [
"node_edges",
"node_inputs"
],
"title": "VueFlowInput",
"type": "object"
}
Fields:
-
node_edges(list[NodeEdge]) -
node_inputs(list[NodeInput]) -
groups(list[FlowfileGroup])
Source code in flowfile_core/flowfile_core/schemas/schemas.py
787 788 789 790 791 792 793 794 795 796 797 798 799 | |
get_global_execution_location()
Calculates the default execution location based on the global settings Returns
ExecutionLocationsLiteral where the current
Source code in flowfile_core/flowfile_core/schemas/schemas.py
79 80 81 82 83 84 85 86 87 88 | |
get_settings_class_for_node_type(node_type, setting_data=None)
Get the settings class for a node type, supporting both standard and user-defined nodes.
setting_data (the raw stored settings dict, when available) lets flows
referencing a custom node that is missing from the store still resolve to
UserDefinedNode instead of failing as an unknown type.
Source code in flowfile_core/flowfile_core/schemas/schemas.py
101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 | |
input_schema
flowfile_core.schemas.input_schema
Classes:
| Name | Description |
|---|---|
ApplyModelSettings |
Settings payload for the Apply Model node. |
CatalogWriteSettings |
Settings for writing data to the catalog. |
DatabaseConnection |
Defines the connection parameters for a database. |
DatabaseSettings |
Defines settings for reading from a database, either via table or query. |
DatabaseWriteSettings |
Defines settings for writing data to a database table. |
EvaluateModelSettings |
Settings payload for the Evaluate Model node. |
ExternalSource |
Base model for data coming from a predefined external source. |
FullDatabaseConnection |
A complete database connection model including the secret password. |
FullDatabaseConnectionInterface |
A database connection model intended for UI display, omitting the password. |
GoogleAnalyticsFilter |
A single filter applied to a GA4 dimension or metric. |
GoogleAnalyticsOrderBy |
A single sort entry applied to the GA4 report. |
GoogleAnalyticsSettings |
UI settings for a Google Analytics 4 reader node. |
InputAvroTable |
Defines settings for reading an Avro file. |
InputCsvTable |
Defines settings for reading a CSV file. |
InputExcelTable |
Defines settings for reading an Excel file. |
InputIpcTable |
Defines settings for reading an Arrow IPC/Feather file. |
InputJsonTable |
Defines settings for reading a JSON file. |
InputNdjsonTable |
Defines settings for reading a newline-delimited JSON file. |
InputParquetTable |
Defines settings for reading a Parquet file. |
InputTableBase |
Base settings for input file operations. |
KafkaSourceSettings |
Configuration for reading from a Kafka/Redpanda topic. |
MinimalFieldInfo |
Represents the most basic information about a data field (column). |
NewDirectory |
Defines the information required to create a new directory. |
NodeApiResponse |
Settings for a node that marks its input as the body of an HTTP API response. |
NodeApplyModel |
Score data using a previously trained model artifact. |
NodeBase |
Base model for all nodes in a FlowGraph. Contains common metadata. |
NodeCatalogReader |
Settings for a node that reads a table from the catalog. |
NodeCatalogWriter |
Settings for a node that writes its input to the catalog. |
NodeCloudStorageReader |
Settings for a node that reads from a cloud storage service (S3, GCS, etc.). |
NodeCloudStorageWriter |
Settings for a node that writes to a cloud storage service. |
NodeConnection |
Represents a connection (edge) between two nodes in the graph. |
NodeCrossJoin |
Settings for a node that performs a cross join. |
NodeDatabaseReader |
Settings for a node that reads from a database. |
NodeDatabaseWriter |
Settings for a node that writes data to a database. |
NodeDatasource |
Base settings for a node that acts as a data source. |
NodeDescription |
A simple model for updating a node's description text. |
NodeDynamicRename |
Settings for a node that renames many columns at once via a single rule. |
NodeEvaluateModel |
Compute model-quality metrics by comparing actual and predicted columns. |
NodeExploreData |
Settings for a node that provides an interactive data exploration interface. |
NodeExternalSource |
Settings for a node that connects to a registered external data source. |
NodeFilter |
Settings for a node that filters rows based on a condition. |
NodeFlowInput |
Named source placeholder inside a subflow. |
NodeFlowOutput |
Named passthrough sink marking a subflow output; multiple allowed per flow. |
NodeFormula |
Settings for a node that applies a formula to create/modify a column. |
NodeFuzzyMatch |
Settings for a node that performs a fuzzy join based on string similarity. |
NodeGoogleAnalyticsReader |
Settings for a node that reads from a Google Analytics 4 property. |
NodeGraphSolver |
Settings for a node that solves graph-based problems (e.g., connected components). |
NodeGroupBy |
Settings for a node that performs a group-by and aggregation operation. |
NodeInputConnection |
Represents the input side of a connection between two nodes. |
NodeJoin |
Settings for a node that performs a standard SQL-style join. |
NodeKafkaSource |
Settings for a node that reads from a Kafka or Redpanda topic. |
NodeManualInput |
Settings for a node that allows direct data entry in the UI. |
NodeMultiInput |
A base model for any node that takes multiple data inputs. |
NodeOutput |
Settings for a node that writes its input to a file. |
NodeOutputConnection |
Represents the output side of a connection between two nodes. |
NodePivot |
Settings for a node that pivots data from a long to a wide format. |
NodePolarsCode |
Settings for a node that executes arbitrary user-provided Polars code. |
NodePromise |
A placeholder node for an operation that has not yet been configured. |
NodePythonScript |
Node that executes Python code on a kernel container. |
NodeRandomSplit |
Settings for a node that randomly partitions rows into N labeled outputs. |
NodeRead |
Settings for a node that reads data from a file. |
NodeRecordCount |
Settings for a node that counts the number of records. |
NodeRecordId |
Settings for a node that adds a unique record ID column. |
NodeRestApiReader |
Settings for a node that reads from a REST API. |
NodeRunFlow |
Settings for a node that executes a catalog-registered flow as a subflow. |
NodeSample |
Settings for a node that samples a subset of the data. |
NodeSelect |
Settings for a node that selects, renames, and reorders columns. |
NodeSingleInput |
A base model for any node that takes a single data input. |
NodeSort |
Settings for a node that sorts the data by one or more columns. |
NodeSqlQuery |
Settings for a node that executes a SQL query against connected data sources. |
NodeTextToRows |
Settings for a node that splits a text column into multiple rows. |
NodeTrainModel |
Train an ML model (regression or classification) and optionally publish it to the catalog. |
NodeUnion |
Settings for a node that concatenates multiple data inputs. |
NodeUnique |
Settings for a node that returns the unique rows from the data. |
NodeUnpivot |
Settings for a node that unpivots data from a wide to a long format. |
NodeWaitFor |
Pass-through node that enforces ordering on extra dependency inputs. |
NodeWindowFunctions |
Settings for a node that adds rolling, cumulative, rank or tile columns. |
NotebookCell |
A single cell in the notebook editor. |
OutputAvroTable |
Defines settings for writing an Avro file. |
OutputCsvTable |
Defines settings for writing a CSV file. |
OutputExcelTable |
Defines settings for writing an Excel file. |
OutputFieldConfig |
Configuration for output field validation and transformation behavior. |
OutputFieldInfo |
Field information with optional default value for output field configuration. |
OutputIpcTable |
Defines settings for writing an Arrow IPC/Feather file. |
OutputNdjsonTable |
Defines settings for writing a newline-delimited JSON file. |
OutputParquetTable |
Defines settings for writing a Parquet file. |
OutputSettings |
Defines the complete settings for an output node. |
PythonScriptInput |
Settings for Python code execution on a kernel. |
RandomSplitGroup |
A single output partition in a random split. |
RawData |
Represents data in a raw, columnar format for manual input. |
ReceivedTable |
Model for defining a table received from an external source. |
RemoveItem |
Represents a single item to be removed from a directory or list. |
RemoveItemsInput |
Defines a list of items to be removed. |
RestApiAuthSettings |
Authentication settings for a REST API reader node. |
RestApiPaginationSettings |
Pagination strategy and parameters for a REST API reader node. |
RestApiSettings |
UI settings for a REST API reader node. |
RunFlowParameterBinding |
How one subflow parameter gets its value for a run_flow execution. |
SampleUsers |
Settings for generating a sample dataset of users. |
Scd2Settings |
Slowly-changing-dimension type 2 configuration for a catalog write. |
SubflowReference |
Reference to a catalog-registered flow. |
TrainModelSettings |
Settings payload for the Train Model node. |
UserDefinedNode |
Settings for a node that contains the user defined node information |
ApplyModelSettings
pydantic-model
Bases: BaseModel
Settings payload for the Apply Model node.
Two model sources are supported:
"upstream"(default): pick a Train Model node from somewhere in this flow's upstream chain. The model file is read from the flow's cache directory using the train node's id — works at design time, no catalog round-trip needed."catalog": fall back to the existing catalog lookup by name/version.
Show JSON schema:
{
"description": "Settings payload for the Apply Model node.\n\nTwo model sources are supported:\n\n- ``\"upstream\"`` (default): pick a Train Model node from somewhere in this\n flow's upstream chain. The model file is read from the flow's cache\n directory using the train node's id \u2014 works at design time, no catalog\n round-trip needed.\n- ``\"catalog\"``: fall back to the existing catalog lookup by name/version.",
"properties": {
"source": {
"default": "upstream",
"enum": [
"upstream",
"catalog"
],
"title": "Source",
"type": "string"
},
"upstream_node_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Upstream Node Id"
},
"model_name": {
"default": "",
"title": "Model Name",
"type": "string"
},
"model_version": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Model Version"
},
"namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Id"
},
"namespace_full_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Full Name"
},
"output_column": {
"default": "prediction",
"title": "Output Column",
"type": "string"
}
},
"title": "ApplyModelSettings",
"type": "object"
}
Config:
protected_namespaces:()
Fields:
-
source(Literal['upstream', 'catalog']) -
upstream_node_id(int | None) -
model_name(str) -
model_version(int | None) -
namespace_id(int | None) -
namespace_full_name(str | None) -
output_column(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 | |
CatalogWriteSettings
pydantic-model
Bases: BaseModel
Settings for writing data to the catalog.
The target namespace is referenced name-first: namespace_full_name ("catalog.schema") is
the portable reference that survives recreation on another machine; namespace_id is a numeric
fallback for flows saved before names were stored.
Show JSON schema:
{
"$defs": {
"Scd2Settings": {
"description": "Slowly-changing-dimension type 2 configuration for a catalog write.\n\nThe business key is ``CatalogWriteSettings.merge_keys`` \u2014 this block only carries the\nchange-detection scope and the names of the four generated columns. It is persisted verbatim\nonto the catalog table record (``CatalogTable.scd2_config``) so a reader can filter history\nwithout ever reading a writer node's settings.",
"properties": {
"compare_columns": {
"items": {
"type": "string"
},
"title": "Compare Columns",
"type": "array"
},
"full_snapshot": {
"default": false,
"title": "Full Snapshot",
"type": "boolean"
},
"partition_on_current": {
"default": true,
"title": "Partition On Current",
"type": "boolean"
},
"surrogate_key_column": {
"default": "sk",
"title": "Surrogate Key Column",
"type": "string"
},
"valid_from_column": {
"default": "valid_from",
"title": "Valid From Column",
"type": "string"
},
"valid_to_column": {
"default": "valid_to",
"title": "Valid To Column",
"type": "string"
},
"is_current_column": {
"default": "is_current",
"title": "Is Current Column",
"type": "string"
}
},
"title": "Scd2Settings",
"type": "object"
}
},
"description": "Settings for writing data to the catalog.\n\nThe target namespace is referenced name-first: ``namespace_full_name`` (``\"catalog.schema\"``) is\nthe portable reference that survives recreation on another machine; ``namespace_id`` is a numeric\nfallback for flows saved before names were stored.",
"properties": {
"table_name": {
"default": "",
"title": "Table Name",
"type": "string"
},
"namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Id"
},
"namespace_full_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Full Name"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"write_mode": {
"default": "overwrite",
"enum": [
"overwrite",
"error",
"append",
"upsert",
"update",
"delete",
"scd2",
"virtual"
],
"title": "Write Mode",
"type": "string"
},
"merge_keys": {
"items": {
"type": "string"
},
"title": "Merge Keys",
"type": "array"
},
"partition_by": {
"items": {
"type": "string"
},
"title": "Partition By",
"type": "array"
},
"scd2": {
"anyOf": [
{
"$ref": "#/$defs/Scd2Settings"
},
{
"type": "null"
}
],
"default": null
}
},
"title": "CatalogWriteSettings",
"type": "object"
}
Fields:
-
table_name(str) -
namespace_id(int | None) -
namespace_full_name(str | None) -
description(str | None) -
write_mode(Literal['overwrite', 'error', 'append', 'upsert', 'update', 'delete', 'scd2', 'virtual']) -
merge_keys(list[str]) -
partition_by(list[str]) -
scd2(Scd2Settings | None)
Validators:
-
_validate_merge_keys
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 | |
DatabaseConnection
pydantic-model
Bases: BaseModel
Defines the connection parameters for a database.
Show JSON schema:
{
"description": "Defines the connection parameters for a database.",
"properties": {
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Username"
},
"password_ref": {
"anyOf": [
{
"description": "An ID referencing an encrypted secret.",
"maxLength": 100,
"minLength": 1,
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Password Ref"
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
}
},
"title": "DatabaseConnection",
"type": "object"
}
Fields:
-
database_type(str) -
username(str | None) -
password_ref(SecretRef | None) -
host(str | None) -
port(int | None) -
database(str | None) -
url(str | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 | |
DatabaseSettings
pydantic-model
Bases: BaseModel
Defines settings for reading from a database, either via table or query.
Show JSON schema:
{
"$defs": {
"DatabaseConnection": {
"description": "Defines the connection parameters for a database.",
"properties": {
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Username"
},
"password_ref": {
"anyOf": [
{
"description": "An ID referencing an encrypted secret.",
"maxLength": 100,
"minLength": 1,
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Password Ref"
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
}
},
"title": "DatabaseConnection",
"type": "object"
}
},
"description": "Defines settings for reading from a database, either via table or query.",
"properties": {
"connection_mode": {
"anyOf": [
{
"enum": [
"inline",
"reference"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "inline",
"title": "Connection Mode"
},
"database_connection": {
"anyOf": [
{
"$ref": "#/$defs/DatabaseConnection"
},
{
"type": "null"
}
],
"default": null
},
"database_connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database Connection Name"
},
"schema_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema Name"
},
"table_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Table Name"
},
"query": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Query"
},
"query_mode": {
"default": "table",
"enum": [
"query",
"table",
"reference"
],
"title": "Query Mode",
"type": "string"
}
},
"title": "DatabaseSettings",
"type": "object"
}
Fields:
-
connection_mode(Literal['inline', 'reference'] | None) -
database_connection(DatabaseConnection | None) -
database_connection_name(str | None) -
schema_name(str | None) -
table_name(str | None) -
query(str | None) -
query_mode(Literal['query', 'table', 'reference'])
Validators:
-
validate_sql_identifier→table_name,schema_name -
validate_table_or_query
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 | |
DatabaseWriteSettings
pydantic-model
Bases: BaseModel
Defines settings for writing data to a database table.
Show JSON schema:
{
"$defs": {
"DatabaseConnection": {
"description": "Defines the connection parameters for a database.",
"properties": {
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Username"
},
"password_ref": {
"anyOf": [
{
"description": "An ID referencing an encrypted secret.",
"maxLength": 100,
"minLength": 1,
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Password Ref"
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
}
},
"title": "DatabaseConnection",
"type": "object"
}
},
"description": "Defines settings for writing data to a database table.",
"properties": {
"connection_mode": {
"anyOf": [
{
"enum": [
"inline",
"reference"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "inline",
"title": "Connection Mode"
},
"database_connection": {
"anyOf": [
{
"$ref": "#/$defs/DatabaseConnection"
},
{
"type": "null"
}
],
"default": null
},
"database_connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database Connection Name"
},
"table_name": {
"title": "Table Name",
"type": "string"
},
"schema_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema Name"
},
"if_exists": {
"anyOf": [
{
"enum": [
"append",
"replace",
"fail"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "append",
"title": "If Exists"
}
},
"required": [
"table_name"
],
"title": "DatabaseWriteSettings",
"type": "object"
}
Fields:
-
connection_mode(Literal['inline', 'reference'] | None) -
database_connection(DatabaseConnection | None) -
database_connection_name(str | None) -
table_name(str) -
schema_name(str | None) -
if_exists(Literal['append', 'replace', 'fail'] | None)
Validators:
-
validate_sql_identifier→table_name,schema_name
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 | |
EvaluateModelSettings
pydantic-model
Bases: BaseModel
Settings payload for the Evaluate Model node.
Decoupled from any specific Train/Apply pair: takes a dataframe that
already contains both the actual target column and a prediction column
and emits a long-form (metric, value) frame. Reusable on training,
test, or hold-out splits.
task_type="auto" resolves the metric set from an upstream Train
Model node when one is configured; otherwise defaults to regression.
Show JSON schema:
{
"description": "Settings payload for the Evaluate Model node.\n\nDecoupled from any specific Train/Apply pair: takes a dataframe that\nalready contains both the actual target column and a prediction column\nand emits a long-form ``(metric, value)`` frame. Reusable on training,\ntest, or hold-out splits.\n\n``task_type=\"auto\"`` resolves the metric set from an upstream Train\nModel node when one is configured; otherwise defaults to ``regression``.",
"properties": {
"actual_column": {
"default": "",
"title": "Actual Column",
"type": "string"
},
"predicted_column": {
"default": "prediction",
"title": "Predicted Column",
"type": "string"
},
"task_type": {
"default": "auto",
"enum": [
"auto",
"regression",
"classification"
],
"title": "Task Type",
"type": "string"
},
"upstream_train_node_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Upstream Train Node Id"
}
},
"title": "EvaluateModelSettings",
"type": "object"
}
Config:
protected_namespaces:()
Fields:
-
actual_column(str) -
predicted_column(str) -
task_type(Literal['auto', 'regression', 'classification']) -
upstream_train_node_id(int | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 | |
ExternalSource
pydantic-model
Bases: BaseModel
Base model for data coming from a predefined external source.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
}
},
"description": "Base model for data coming from a predefined external source.",
"properties": {
"orientation": {
"default": "row",
"title": "Orientation",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
}
},
"title": "ExternalSource",
"type": "object"
}
Fields:
-
orientation(str) -
fields(list[MinimalFieldInfo] | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1135 1136 1137 1138 1139 | |
FullDatabaseConnection
pydantic-model
Bases: BaseModel
A complete database connection model including the secret password.
Show JSON schema:
{
"description": "A complete database connection model including the secret password.",
"properties": {
"connection_name": {
"title": "Connection Name",
"type": "string"
},
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"title": "Username",
"type": "string"
},
"password": {
"format": "password",
"title": "Password",
"type": "string",
"writeOnly": true
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"ssl_enabled": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Ssl Enabled"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
}
},
"required": [
"connection_name",
"username",
"password"
],
"title": "FullDatabaseConnection",
"type": "object"
}
Fields:
-
connection_name(str) -
database_type(str) -
username(str) -
password(SecretStr) -
host(str | None) -
port(int | None) -
database(str | None) -
ssl_enabled(bool | None) -
url(str | None)
Validators:
-
normalize_database_type→database_type
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 | |
FullDatabaseConnectionInterface
pydantic-model
Bases: BaseModel
A database connection model intended for UI display, omitting the password.
Show JSON schema:
{
"$defs": {
"AccessInfo": {
"description": "How the requesting user can access a resource; attached to list/detail responses.",
"properties": {
"is_owner": {
"title": "Is Owner",
"type": "boolean"
},
"access_level": {
"enum": [
"owner",
"manage",
"use"
],
"title": "Access Level",
"type": "string"
},
"shared_by": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Shared By"
}
},
"required": [
"is_owner",
"access_level"
],
"title": "AccessInfo",
"type": "object"
}
},
"description": "A database connection model intended for UI display, omitting the password.",
"properties": {
"connection_name": {
"title": "Connection Name",
"type": "string"
},
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"title": "Username",
"type": "string"
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"ssl_enabled": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Ssl Enabled"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
},
"id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Id"
},
"access": {
"anyOf": [
{
"$ref": "#/$defs/AccessInfo"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"connection_name",
"username"
],
"title": "FullDatabaseConnectionInterface",
"type": "object"
}
Fields:
-
connection_name(str) -
database_type(str) -
username(str) -
host(str | None) -
port(int | None) -
database(str | None) -
ssl_enabled(bool | None) -
url(str | None) -
id(int | None) -
access(AccessInfo | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 | |
GoogleAnalyticsFilter
pydantic-model
Bases: BaseModel
A single filter applied to a GA4 dimension or metric.
field must match one of the selected dimensions or metrics; the worker
auto-routes the filter into either the request's dimension_filter (for
string-typed dimensions) or metric_filter (for numeric-typed metrics).
Supported operators — strings (dimensions):
- equals, not_equals
- contains, begins_with, ends_with
- regex (full regex match)
- in_list, not_in_list (comma-separated value)
Supported operators — numeric (metrics):
- equals, not_equals
- less_than, less_equal, greater_than, greater_equal
- between (comma-separated "low,high")
Multiple filters on the same kind are AND-combined. String matching is
case-insensitive by default (case_sensitive=False below).
Show JSON schema:
{
"description": "A single filter applied to a GA4 dimension or metric.\n\n``field`` must match one of the selected dimensions or metrics; the worker\nauto-routes the filter into either the request's ``dimension_filter`` (for\nstring-typed dimensions) or ``metric_filter`` (for numeric-typed metrics).\n\nSupported operators \u2014 strings (dimensions):\n - ``equals``, ``not_equals``\n - ``contains``, ``begins_with``, ``ends_with``\n - ``regex`` (full regex match)\n - ``in_list``, ``not_in_list`` (comma-separated ``value``)\n\nSupported operators \u2014 numeric (metrics):\n - ``equals``, ``not_equals``\n - ``less_than``, ``less_equal``, ``greater_than``, ``greater_equal``\n - ``between`` (comma-separated ``\"low,high\"``)\n\nMultiple filters on the same kind are AND-combined. String matching is\ncase-insensitive by default (``case_sensitive=False`` below).",
"properties": {
"field": {
"title": "Field",
"type": "string"
},
"operator": {
"title": "Operator",
"type": "string"
},
"value": {
"default": "",
"title": "Value",
"type": "string"
},
"case_sensitive": {
"default": false,
"title": "Case Sensitive",
"type": "boolean"
}
},
"required": [
"field",
"operator"
],
"title": "GoogleAnalyticsFilter",
"type": "object"
}
Fields:
-
field(str) -
operator(str) -
value(str) -
case_sensitive(bool)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 | |
GoogleAnalyticsOrderBy
pydantic-model
Bases: BaseModel
A single sort entry applied to the GA4 report.
field must match one of the selected dimensions or metrics; the worker
routes it into a DimensionOrderBy or MetricOrderBy accordingly.
descending=True produces a descending sort. Sort entries are applied in
list order.
Show JSON schema:
{
"description": "A single sort entry applied to the GA4 report.\n\n``field`` must match one of the selected dimensions or metrics; the worker\nroutes it into a ``DimensionOrderBy`` or ``MetricOrderBy`` accordingly.\n``descending=True`` produces a descending sort. Sort entries are applied in\nlist order.",
"properties": {
"field": {
"title": "Field",
"type": "string"
},
"descending": {
"default": false,
"title": "Descending",
"type": "boolean"
}
},
"required": [
"field"
],
"title": "GoogleAnalyticsOrderBy",
"type": "object"
}
Fields:
-
field(str) -
descending(bool)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 | |
GoogleAnalyticsSettings
pydantic-model
Bases: BaseModel
UI settings for a Google Analytics 4 reader node.
Credentials are NOT stored inline: ga_connection_name is a reference to
a Google Analytics connection managed under /ga_connections (whose
service-account JSON is encrypted at rest).
Show JSON schema:
{
"$defs": {
"GoogleAnalyticsFilter": {
"description": "A single filter applied to a GA4 dimension or metric.\n\n``field`` must match one of the selected dimensions or metrics; the worker\nauto-routes the filter into either the request's ``dimension_filter`` (for\nstring-typed dimensions) or ``metric_filter`` (for numeric-typed metrics).\n\nSupported operators \u2014 strings (dimensions):\n - ``equals``, ``not_equals``\n - ``contains``, ``begins_with``, ``ends_with``\n - ``regex`` (full regex match)\n - ``in_list``, ``not_in_list`` (comma-separated ``value``)\n\nSupported operators \u2014 numeric (metrics):\n - ``equals``, ``not_equals``\n - ``less_than``, ``less_equal``, ``greater_than``, ``greater_equal``\n - ``between`` (comma-separated ``\"low,high\"``)\n\nMultiple filters on the same kind are AND-combined. String matching is\ncase-insensitive by default (``case_sensitive=False`` below).",
"properties": {
"field": {
"title": "Field",
"type": "string"
},
"operator": {
"title": "Operator",
"type": "string"
},
"value": {
"default": "",
"title": "Value",
"type": "string"
},
"case_sensitive": {
"default": false,
"title": "Case Sensitive",
"type": "boolean"
}
},
"required": [
"field",
"operator"
],
"title": "GoogleAnalyticsFilter",
"type": "object"
},
"GoogleAnalyticsOrderBy": {
"description": "A single sort entry applied to the GA4 report.\n\n``field`` must match one of the selected dimensions or metrics; the worker\nroutes it into a ``DimensionOrderBy`` or ``MetricOrderBy`` accordingly.\n``descending=True`` produces a descending sort. Sort entries are applied in\nlist order.",
"properties": {
"field": {
"title": "Field",
"type": "string"
},
"descending": {
"default": false,
"title": "Descending",
"type": "boolean"
}
},
"required": [
"field"
],
"title": "GoogleAnalyticsOrderBy",
"type": "object"
}
},
"description": "UI settings for a Google Analytics 4 reader node.\n\nCredentials are NOT stored inline: ``ga_connection_name`` is a reference to\na Google Analytics connection managed under ``/ga_connections`` (whose\nservice-account JSON is encrypted at rest).",
"properties": {
"ga_connection_name": {
"title": "Ga Connection Name",
"type": "string"
},
"property_id": {
"title": "Property Id",
"type": "string"
},
"start_date": {
"default": "7daysAgo",
"title": "Start Date",
"type": "string"
},
"end_date": {
"default": "yesterday",
"title": "End Date",
"type": "string"
},
"metrics": {
"items": {
"type": "string"
},
"title": "Metrics",
"type": "array"
},
"dimensions": {
"items": {
"type": "string"
},
"title": "Dimensions",
"type": "array"
},
"limit": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Limit"
},
"filters": {
"items": {
"$ref": "#/$defs/GoogleAnalyticsFilter"
},
"title": "Filters",
"type": "array"
},
"order_bys": {
"items": {
"$ref": "#/$defs/GoogleAnalyticsOrderBy"
},
"title": "Order Bys",
"type": "array"
}
},
"required": [
"ga_connection_name",
"property_id"
],
"title": "GoogleAnalyticsSettings",
"type": "object"
}
Fields:
-
ga_connection_name(str) -
property_id(str) -
start_date(str) -
end_date(str) -
metrics(list[str]) -
dimensions(list[str]) -
limit(int | None) -
filters(list[GoogleAnalyticsFilter]) -
order_bys(list[GoogleAnalyticsOrderBy])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 | |
InputAvroTable
pydantic-model
Bases: InputTableBase
Defines settings for reading an Avro file.
Show JSON schema:
{
"description": "Defines settings for reading an Avro file.",
"properties": {
"file_type": {
"const": "avro",
"default": "avro",
"title": "File Type",
"type": "string"
}
},
"title": "InputAvroTable",
"type": "object"
}
Fields:
-
file_type(Literal['avro'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
184 185 186 187 | |
InputCsvTable
pydantic-model
Bases: InputTableBase
Defines settings for reading a CSV file.
Show JSON schema:
{
"description": "Defines settings for reading a CSV file.",
"properties": {
"file_type": {
"const": "csv",
"default": "csv",
"title": "File Type",
"type": "string"
},
"reference": {
"default": "",
"title": "Reference",
"type": "string"
},
"starting_from_line": {
"default": 0,
"title": "Starting From Line",
"type": "integer"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
},
"parquet_ref": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Parquet Ref"
},
"row_delimiter": {
"default": "\n",
"title": "Row Delimiter",
"type": "string"
},
"quote_char": {
"default": "\"",
"title": "Quote Char",
"type": "string"
},
"infer_schema_length": {
"default": 10000,
"title": "Infer Schema Length",
"type": "integer"
},
"infer_schema": {
"default": true,
"title": "Infer Schema",
"type": "boolean"
},
"truncate_ragged_lines": {
"default": false,
"title": "Truncate Ragged Lines",
"type": "boolean"
},
"ignore_errors": {
"default": false,
"title": "Ignore Errors",
"type": "boolean"
}
},
"title": "InputCsvTable",
"type": "object"
}
Fields:
-
file_type(Literal['csv']) -
reference(str) -
starting_from_line(int) -
delimiter(str) -
has_headers(bool) -
encoding(str) -
parquet_ref(str | None) -
row_delimiter(str) -
quote_char(str) -
infer_schema_length(int) -
infer_schema(bool) -
truncate_ragged_lines(bool) -
ignore_errors(bool)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 | |
InputExcelTable
pydantic-model
Bases: InputTableBase
Defines settings for reading an Excel file.
Show JSON schema:
{
"description": "Defines settings for reading an Excel file.",
"properties": {
"file_type": {
"const": "excel",
"default": "excel",
"title": "File Type",
"type": "string"
},
"sheet_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Sheet Name"
},
"start_row": {
"default": 0,
"title": "Start Row",
"type": "integer"
},
"start_column": {
"default": 0,
"title": "Start Column",
"type": "integer"
},
"end_row": {
"default": 0,
"title": "End Row",
"type": "integer"
},
"end_column": {
"default": 0,
"title": "End Column",
"type": "integer"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"type_inference": {
"default": false,
"title": "Type Inference",
"type": "boolean"
}
},
"title": "InputExcelTable",
"type": "object"
}
Fields:
-
file_type(Literal['excel']) -
sheet_name(str | None) -
start_row(int) -
start_column(int) -
end_row(int) -
end_column(int) -
has_headers(bool) -
type_inference(bool)
Validators:
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 | |
validate_range_values()
pydantic-validator
Validates that the Excel cell range is logical.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
159 160 161 162 163 164 165 166 167 168 169 | |
InputIpcTable
pydantic-model
Bases: InputTableBase
Defines settings for reading an Arrow IPC/Feather file.
Show JSON schema:
{
"description": "Defines settings for reading an Arrow IPC/Feather file.",
"properties": {
"file_type": {
"const": "ipc",
"default": "ipc",
"title": "File Type",
"type": "string"
}
},
"title": "InputIpcTable",
"type": "object"
}
Fields:
-
file_type(Literal['ipc'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
172 173 174 175 | |
InputJsonTable
pydantic-model
Bases: InputCsvTable
Defines settings for reading a JSON file.
Show JSON schema:
{
"description": "Defines settings for reading a JSON file.",
"properties": {
"file_type": {
"const": "json",
"default": "json",
"title": "File Type",
"type": "string"
},
"reference": {
"default": "",
"title": "Reference",
"type": "string"
},
"starting_from_line": {
"default": 0,
"title": "Starting From Line",
"type": "integer"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
},
"parquet_ref": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Parquet Ref"
},
"row_delimiter": {
"default": "\n",
"title": "Row Delimiter",
"type": "string"
},
"quote_char": {
"default": "\"",
"title": "Quote Char",
"type": "string"
},
"infer_schema_length": {
"default": 10000,
"title": "Infer Schema Length",
"type": "integer"
},
"infer_schema": {
"default": true,
"title": "Infer Schema",
"type": "boolean"
},
"truncate_ragged_lines": {
"default": false,
"title": "Truncate Ragged Lines",
"type": "boolean"
},
"ignore_errors": {
"default": false,
"title": "Ignore Errors",
"type": "boolean"
}
},
"title": "InputJsonTable",
"type": "object"
}
Fields:
-
reference(str) -
starting_from_line(int) -
delimiter(str) -
has_headers(bool) -
encoding(str) -
parquet_ref(str | None) -
row_delimiter(str) -
quote_char(str) -
infer_schema_length(int) -
infer_schema(bool) -
truncate_ragged_lines(bool) -
ignore_errors(bool) -
file_type(Literal['json'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
135 136 137 138 | |
InputNdjsonTable
pydantic-model
Bases: InputTableBase
Defines settings for reading a newline-delimited JSON file.
Show JSON schema:
{
"description": "Defines settings for reading a newline-delimited JSON file.",
"properties": {
"file_type": {
"const": "ndjson",
"default": "ndjson",
"title": "File Type",
"type": "string"
}
},
"title": "InputNdjsonTable",
"type": "object"
}
Fields:
-
file_type(Literal['ndjson'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
178 179 180 181 | |
InputParquetTable
pydantic-model
Bases: InputTableBase
Defines settings for reading a Parquet file.
Show JSON schema:
{
"description": "Defines settings for reading a Parquet file.",
"properties": {
"file_type": {
"const": "parquet",
"default": "parquet",
"title": "File Type",
"type": "string"
}
},
"title": "InputParquetTable",
"type": "object"
}
Fields:
-
file_type(Literal['parquet'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
141 142 143 144 | |
InputTableBase
pydantic-model
Bases: BaseModel
Base settings for input file operations.
Show JSON schema:
{
"description": "Base settings for input file operations.",
"properties": {
"file_type": {
"title": "File Type",
"type": "string"
}
},
"required": [
"file_type"
],
"title": "InputTableBase",
"type": "object"
}
Fields:
-
file_type(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
111 112 113 114 | |
KafkaSourceSettings
pydantic-model
Bases: BaseModel
Configuration for reading from a Kafka/Redpanda topic.
Show JSON schema:
{
"description": "Configuration for reading from a Kafka/Redpanda topic.",
"properties": {
"kafka_connection_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Kafka Connection Id"
},
"kafka_connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Kafka Connection Name"
},
"topic_name": {
"default": "",
"title": "Topic Name",
"type": "string"
},
"value_format": {
"const": "json",
"default": "json",
"title": "Value Format",
"type": "string"
},
"sync_name": {
"default": "",
"title": "Sync Name",
"type": "string"
},
"start_offset": {
"default": "latest",
"enum": [
"earliest",
"latest"
],
"title": "Start Offset",
"type": "string"
},
"max_messages": {
"default": 100000,
"title": "Max Messages",
"type": "integer"
},
"poll_timeout_seconds": {
"default": 30.0,
"title": "Poll Timeout Seconds",
"type": "number"
}
},
"title": "KafkaSourceSettings",
"type": "object"
}
Fields:
-
kafka_connection_id(int | None) -
kafka_connection_name(str | None) -
topic_name(str) -
value_format(Literal['json']) -
sync_name(str) -
start_offset(Literal['earliest', 'latest']) -
max_messages(int) -
poll_timeout_seconds(float)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 | |
MinimalFieldInfo
pydantic-model
Bases: BaseModel
Represents the most basic information about a data field (column).
Show JSON schema:
{
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
}
Fields:
-
name(str) -
data_type(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
82 83 84 85 86 | |
NewDirectory
pydantic-model
Bases: BaseModel
Defines the information required to create a new directory.
Show JSON schema:
{
"description": "Defines the information required to create a new directory.",
"properties": {
"source_path": {
"title": "Source Path",
"type": "string"
},
"dir_name": {
"title": "Dir Name",
"type": "string"
}
},
"required": [
"source_path",
"dir_name"
],
"title": "NewDirectory",
"type": "object"
}
Fields:
-
source_path(str) -
dir_name(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
61 62 63 64 65 | |
NodeApiResponse
pydantic-model
Bases: NodeSingleInput
Settings for a node that marks its input as the body of an HTTP API response.
This node is a sink (one input, no output). When the flow is published as an API endpoint, the data flowing into this node is serialized and returned to the caller. During interactive runs it is a pass-through (its result equals its input), so previews keep working.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that marks its input as the body of an HTTP API response.\n\nThis node is a sink (one input, no output). When the flow is published as an\nAPI endpoint, the data flowing into this node is serialized and returned to the\ncaller. During interactive runs it is a pass-through (its result equals its\ninput), so previews keep working.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"orientation": {
"default": "records",
"enum": [
"records",
"columns"
],
"title": "Orientation",
"type": "string"
},
"max_rows": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Rows"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeApiResponse",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
orientation(Literal['records', 'columns']) -
max_rows(int | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 | |
get_default_description()
Describes the API response shape.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1516 1517 1518 1519 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeApplyModel
pydantic-model
Bases: NodeSingleInput
Score data using a previously trained model artifact.
Show JSON schema:
{
"$defs": {
"ApplyModelSettings": {
"description": "Settings payload for the Apply Model node.\n\nTwo model sources are supported:\n\n- ``\"upstream\"`` (default): pick a Train Model node from somewhere in this\n flow's upstream chain. The model file is read from the flow's cache\n directory using the train node's id \u2014 works at design time, no catalog\n round-trip needed.\n- ``\"catalog\"``: fall back to the existing catalog lookup by name/version.",
"properties": {
"source": {
"default": "upstream",
"enum": [
"upstream",
"catalog"
],
"title": "Source",
"type": "string"
},
"upstream_node_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Upstream Node Id"
},
"model_name": {
"default": "",
"title": "Model Name",
"type": "string"
},
"model_version": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Model Version"
},
"namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Id"
},
"namespace_full_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Full Name"
},
"output_column": {
"default": "prediction",
"title": "Output Column",
"type": "string"
}
},
"title": "ApplyModelSettings",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Score data using a previously trained model artifact.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"apply_input": {
"$ref": "#/$defs/ApplyModelSettings"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeApplyModel",
"type": "object"
}
Config:
protected_namespaces:()
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
apply_input(ApplyModelSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeBase
pydantic-model
Bases: BaseModel
Base model for all nodes in a FlowGraph. Contains common metadata.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Base model for all nodes in a FlowGraph. Contains common metadata.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeBase",
"type": "object"
}
Config:
arbitrary_types_allowed:True
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeCatalogReader
pydantic-model
Bases: NodeBase
Settings for a node that reads a table from the catalog.
Resolution priority at runtime: catalog_table_id > catalog_full_table_name >
(catalog_table_name, catalog_namespace_id). The qualified form
(catalog_full_table_name = "schema.table") is the preferred human-facing
identifier when an id isn't available.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that reads a table from the catalog.\n\nResolution priority at runtime: ``catalog_table_id`` > ``catalog_full_table_name`` >\n``(catalog_table_name, catalog_namespace_id)``. The qualified form\n(``catalog_full_table_name`` = ``\"schema.table\"``) is the preferred human-facing\nidentifier when an id isn't available.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"catalog_table_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Catalog Table Id"
},
"catalog_full_table_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Catalog Full Table Name"
},
"catalog_table_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Catalog Table Name"
},
"catalog_namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Catalog Namespace Id"
},
"delta_version": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Delta Version"
},
"scd2_view": {
"anyOf": [
{
"enum": [
"active",
"all",
"active_at"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Scd2 View"
},
"scd2_as_of": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Scd2 As Of"
},
"sql_query": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Sql Query"
},
"is_virtual_optimized": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Is Virtual Optimized"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeCatalogReader",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
catalog_table_id(int | None) -
catalog_full_table_name(str | None) -
catalog_table_name(str | None) -
catalog_namespace_id(int | None) -
delta_version(int | None) -
scd2_view(Literal['active', 'all', 'active_at'] | None) -
scd2_as_of(str | None) -
sql_query(str | None) -
is_virtual_optimized(bool | None)
Validators:
-
validate_node_reference→node_reference -
_validate_scd2_view
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeCatalogWriter
pydantic-model
Bases: NodeSingleInput
Settings for a node that writes its input to the catalog.
Show JSON schema:
{
"$defs": {
"CatalogWriteSettings": {
"description": "Settings for writing data to the catalog.\n\nThe target namespace is referenced name-first: ``namespace_full_name`` (``\"catalog.schema\"``) is\nthe portable reference that survives recreation on another machine; ``namespace_id`` is a numeric\nfallback for flows saved before names were stored.",
"properties": {
"table_name": {
"default": "",
"title": "Table Name",
"type": "string"
},
"namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Id"
},
"namespace_full_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Full Name"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Description"
},
"write_mode": {
"default": "overwrite",
"enum": [
"overwrite",
"error",
"append",
"upsert",
"update",
"delete",
"scd2",
"virtual"
],
"title": "Write Mode",
"type": "string"
},
"merge_keys": {
"items": {
"type": "string"
},
"title": "Merge Keys",
"type": "array"
},
"partition_by": {
"items": {
"type": "string"
},
"title": "Partition By",
"type": "array"
},
"scd2": {
"anyOf": [
{
"$ref": "#/$defs/Scd2Settings"
},
{
"type": "null"
}
],
"default": null
}
},
"title": "CatalogWriteSettings",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"Scd2Settings": {
"description": "Slowly-changing-dimension type 2 configuration for a catalog write.\n\nThe business key is ``CatalogWriteSettings.merge_keys`` \u2014 this block only carries the\nchange-detection scope and the names of the four generated columns. It is persisted verbatim\nonto the catalog table record (``CatalogTable.scd2_config``) so a reader can filter history\nwithout ever reading a writer node's settings.",
"properties": {
"compare_columns": {
"items": {
"type": "string"
},
"title": "Compare Columns",
"type": "array"
},
"full_snapshot": {
"default": false,
"title": "Full Snapshot",
"type": "boolean"
},
"partition_on_current": {
"default": true,
"title": "Partition On Current",
"type": "boolean"
},
"surrogate_key_column": {
"default": "sk",
"title": "Surrogate Key Column",
"type": "string"
},
"valid_from_column": {
"default": "valid_from",
"title": "Valid From Column",
"type": "string"
},
"valid_to_column": {
"default": "valid_to",
"title": "Valid To Column",
"type": "string"
},
"is_current_column": {
"default": "is_current",
"title": "Is Current Column",
"type": "string"
}
},
"title": "Scd2Settings",
"type": "object"
}
},
"description": "Settings for a node that writes its input to the catalog.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"catalog_write_settings": {
"$ref": "#/$defs/CatalogWriteSettings"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeCatalogWriter",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
catalog_write_settings(CatalogWriteSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1735 1736 1737 1738 1739 1740 1741 1742 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeCloudStorageReader
pydantic-model
Bases: NodeBase
Settings for a node that reads from a cloud storage service (S3, GCS, etc.).
Show JSON schema:
{
"$defs": {
"CloudStorageReadSettings": {
"description": "Settings for reading from cloud storage",
"properties": {
"auth_mode": {
"default": "auto",
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account",
"auto"
],
"title": "Auth Mode",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Connection Name"
},
"resource_path": {
"title": "Resource Path",
"type": "string"
},
"scan_mode": {
"default": "single_file",
"enum": [
"single_file",
"directory"
],
"title": "Scan Mode",
"type": "string"
},
"file_format": {
"default": "parquet",
"enum": [
"csv",
"parquet",
"json",
"delta",
"iceberg"
],
"title": "File Format",
"type": "string"
},
"csv_has_header": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Csv Has Header"
},
"csv_delimiter": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": ",",
"title": "Csv Delimiter"
},
"csv_encoding": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "utf8",
"title": "Csv Encoding"
},
"delta_version": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Delta Version"
}
},
"required": [
"resource_path"
],
"title": "CloudStorageReadSettings",
"type": "object"
},
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that reads from a cloud storage service (S3, GCS, etc.).",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"cloud_storage_settings": {
"$ref": "#/$defs/CloudStorageReadSettings"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
}
},
"required": [
"flow_id",
"node_id",
"cloud_storage_settings"
],
"title": "NodeCloudStorageReader",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
cloud_storage_settings(CloudStorageReadSettings) -
fields(list[MinimalFieldInfo] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 | |
get_default_description()
Describes the cloud storage source.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1118 1119 1120 1121 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeCloudStorageWriter
pydantic-model
Bases: NodeSingleInput
Settings for a node that writes to a cloud storage service.
Show JSON schema:
{
"$defs": {
"CloudStorageWriteSettings": {
"description": "Settings for writing to cloud storage",
"properties": {
"resource_path": {
"title": "Resource Path",
"type": "string"
},
"write_mode": {
"default": "overwrite",
"enum": [
"overwrite",
"append"
],
"title": "Write Mode",
"type": "string"
},
"file_format": {
"default": "parquet",
"enum": [
"csv",
"parquet",
"json",
"delta"
],
"title": "File Format",
"type": "string"
},
"parquet_compression": {
"default": "snappy",
"enum": [
"snappy",
"gzip",
"brotli",
"lz4",
"zstd"
],
"title": "Parquet Compression",
"type": "string"
},
"csv_delimiter": {
"default": ",",
"title": "Csv Delimiter",
"type": "string"
},
"csv_encoding": {
"default": "utf8",
"title": "Csv Encoding",
"type": "string"
},
"partition_by": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Partition By"
},
"auth_mode": {
"default": "auto",
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account",
"auto"
],
"title": "Auth Mode",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Connection Name"
}
},
"required": [
"resource_path"
],
"title": "CloudStorageWriteSettings",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that writes to a cloud storage service.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"cloud_storage_settings": {
"$ref": "#/$defs/CloudStorageWriteSettings"
}
},
"required": [
"flow_id",
"node_id",
"cloud_storage_settings"
],
"title": "NodeCloudStorageWriter",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
cloud_storage_settings(CloudStorageWriteSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1124 1125 1126 1127 1128 1129 1130 1131 1132 | |
get_default_description()
Describes the cloud storage write target.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1129 1130 1131 1132 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeConnection
pydantic-model
Bases: BaseModel
Represents a connection (edge) between two nodes in the graph.
Show JSON schema:
{
"$defs": {
"NodeInputConnection": {
"description": "Represents the input side of a connection between two nodes.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"connection_class": {
"enum": [
"input-0",
"input-1",
"input-2",
"input-3",
"input-4",
"input-5",
"input-6",
"input-7",
"input-8",
"input-9"
],
"title": "Connection Class",
"type": "string"
}
},
"required": [
"node_id",
"connection_class"
],
"title": "NodeInputConnection",
"type": "object"
},
"NodeOutputConnection": {
"description": "Represents the output side of a connection between two nodes.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"connection_class": {
"enum": [
"output-0",
"output-1",
"output-2",
"output-3",
"output-4",
"output-5",
"output-6",
"output-7",
"output-8",
"output-9"
],
"title": "Connection Class",
"type": "string"
}
},
"required": [
"node_id",
"connection_class"
],
"title": "NodeOutputConnection",
"type": "object"
}
},
"description": "Represents a connection (edge) between two nodes in the graph.",
"properties": {
"input_connection": {
"$ref": "#/$defs/NodeInputConnection"
},
"output_connection": {
"$ref": "#/$defs/NodeOutputConnection"
}
},
"required": [
"input_connection",
"output_connection"
],
"title": "NodeConnection",
"type": "object"
}
Fields:
-
input_connection(NodeInputConnection) -
output_connection(NodeOutputConnection)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 | |
create_from_simple_input(from_id, to_id, input_type='input-0', output_handle='output-0')
classmethod
Creates a standard connection between two nodes.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1834 1835 1836 1837 1838 1839 1840 1841 1842 1843 1844 1845 1846 1847 1848 1849 1850 1851 1852 1853 1854 | |
NodeCrossJoin
pydantic-model
Bases: NodeMultiInput
Settings for a node that performs a cross join.
Show JSON schema:
{
"$defs": {
"CrossJoinInput": {
"description": "Data model for cross join operations.",
"properties": {
"left_select": {
"$ref": "#/$defs/JoinInputs"
},
"right_select": {
"$ref": "#/$defs/JoinInputs"
}
},
"required": [
"left_select",
"right_select"
],
"title": "CrossJoinInput",
"type": "object"
},
"JoinInputs": {
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Settings for a node that performs a cross join.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"auto_generate_selection": {
"default": true,
"title": "Auto Generate Selection",
"type": "boolean"
},
"verify_integrity": {
"default": true,
"title": "Verify Integrity",
"type": "boolean"
},
"cross_join_input": {
"$ref": "#/$defs/CrossJoinInput"
},
"auto_keep_all": {
"default": true,
"title": "Auto Keep All",
"type": "boolean"
},
"auto_keep_right": {
"default": true,
"title": "Auto Keep Right",
"type": "boolean"
},
"auto_keep_left": {
"default": true,
"title": "Auto Keep Left",
"type": "boolean"
}
},
"required": [
"flow_id",
"node_id",
"cross_join_input"
],
"title": "NodeCrossJoin",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
auto_generate_selection(bool) -
verify_integrity(bool) -
cross_join_input(CrossJoinInput) -
auto_keep_all(bool) -
auto_keep_right(bool) -
auto_keep_left(bool)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 | |
get_default_description()
Describes the cross join.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
773 774 775 | |
to_yaml_dict()
Converts the cross join node settings to a dictionary for YAML serialization.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeDatabaseReader
pydantic-model
Bases: NodeBase
Settings for a node that reads from a database.
Show JSON schema:
{
"$defs": {
"DatabaseConnection": {
"description": "Defines the connection parameters for a database.",
"properties": {
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Username"
},
"password_ref": {
"anyOf": [
{
"description": "An ID referencing an encrypted secret.",
"maxLength": 100,
"minLength": 1,
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Password Ref"
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
}
},
"title": "DatabaseConnection",
"type": "object"
},
"DatabaseSettings": {
"description": "Defines settings for reading from a database, either via table or query.",
"properties": {
"connection_mode": {
"anyOf": [
{
"enum": [
"inline",
"reference"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "inline",
"title": "Connection Mode"
},
"database_connection": {
"anyOf": [
{
"$ref": "#/$defs/DatabaseConnection"
},
{
"type": "null"
}
],
"default": null
},
"database_connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database Connection Name"
},
"schema_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema Name"
},
"table_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Table Name"
},
"query": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Query"
},
"query_mode": {
"default": "table",
"enum": [
"query",
"table",
"reference"
],
"title": "Query Mode",
"type": "string"
}
},
"title": "DatabaseSettings",
"type": "object"
},
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that reads from a database.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"database_settings": {
"$ref": "#/$defs/DatabaseSettings"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
}
},
"required": [
"flow_id",
"node_id",
"database_settings"
],
"title": "NodeDatabaseReader",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
database_settings(DatabaseSettings) -
fields(list[MinimalFieldInfo] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 | |
get_default_description()
Describes the database source.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeDatabaseWriter
pydantic-model
Bases: NodeSingleInput
Settings for a node that writes data to a database.
Show JSON schema:
{
"$defs": {
"DatabaseConnection": {
"description": "Defines the connection parameters for a database.",
"properties": {
"database_type": {
"default": "postgresql",
"title": "Database Type",
"type": "string"
},
"username": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Username"
},
"password_ref": {
"anyOf": [
{
"description": "An ID referencing an encrypted secret.",
"maxLength": 100,
"minLength": 1,
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Password Ref"
},
"host": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Host"
},
"port": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Port"
},
"database": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database"
},
"url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Url"
}
},
"title": "DatabaseConnection",
"type": "object"
},
"DatabaseWriteSettings": {
"description": "Defines settings for writing data to a database table.",
"properties": {
"connection_mode": {
"anyOf": [
{
"enum": [
"inline",
"reference"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "inline",
"title": "Connection Mode"
},
"database_connection": {
"anyOf": [
{
"$ref": "#/$defs/DatabaseConnection"
},
{
"type": "null"
}
],
"default": null
},
"database_connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Database Connection Name"
},
"table_name": {
"title": "Table Name",
"type": "string"
},
"schema_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema Name"
},
"if_exists": {
"anyOf": [
{
"enum": [
"append",
"replace",
"fail"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "append",
"title": "If Exists"
}
},
"required": [
"table_name"
],
"title": "DatabaseWriteSettings",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that writes data to a database.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"database_write_settings": {
"$ref": "#/$defs/DatabaseWriteSettings"
}
},
"required": [
"flow_id",
"node_id",
"database_write_settings"
],
"title": "NodeDatabaseWriter",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
database_write_settings(DatabaseWriteSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 | |
get_default_description()
Describes the database write target.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1105 1106 1107 1108 1109 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeDatasource
pydantic-model
Bases: NodeBase
Base settings for a node that acts as a data source.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Base settings for a node that acts as a data source.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"file_ref": {
"default": null,
"title": "File Ref",
"type": "string"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeDatasource",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
file_ref(str)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
853 854 855 856 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeDescription
pydantic-model
Bases: BaseModel
A simple model for updating a node's description text.
Show JSON schema:
{
"description": "A simple model for updating a node's description text.",
"properties": {
"description": {
"default": "",
"title": "Description",
"type": "string"
}
},
"title": "NodeDescription",
"type": "object"
}
Fields:
-
description(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1857 1858 1859 1860 | |
NodeDynamicRename
pydantic-model
Bases: NodeSingleInput
Settings for a node that renames many columns at once via a single rule.
Show JSON schema:
{
"$defs": {
"DynamicRenameInput": {
"description": "Defines settings for a dynamic rename operation.\n\nApplies a single rule (prefix / suffix / formula / first_row) to a set of selected\ncolumns, rather than requiring the user to rename columns one-by-one.\n\nIn formula mode, the flowfile formula syntax is evaluated with `[column_name]`\nbound to each target column's current name; for example `uppercase([column_name])`\nor `\"v2_\" + [column_name]`.\n\nIn first_row mode, the first row of the incoming table is promoted to column\nheaders and then dropped from the data. Non-string values are coerced to `str`;\nnull or empty values raise an error. Selection filters still apply \u2014 only selected\ncolumns are renamed, but the first row is always dropped.",
"properties": {
"rename_mode": {
"default": "prefix",
"enum": [
"prefix",
"suffix",
"formula",
"first_row"
],
"title": "Rename Mode",
"type": "string"
},
"prefix": {
"default": "",
"title": "Prefix",
"type": "string"
},
"suffix": {
"default": "",
"title": "Suffix",
"type": "string"
},
"formula": {
"default": "",
"expression": true,
"title": "Formula",
"type": "string"
},
"selection_mode": {
"default": "all",
"enum": [
"all",
"list",
"data_type"
],
"title": "Selection Mode",
"type": "string"
},
"selected_columns": {
"items": {
"type": "string"
},
"title": "Selected Columns",
"type": "array"
},
"selected_data_type": {
"anyOf": [
{
"enum": [
"Numeric",
"String",
"Date",
"Other",
"Boolean",
"Binary",
"Complex"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Selected Data Type"
}
},
"title": "DynamicRenameInput",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that renames many columns at once via a single rule.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"dynamic_rename_input": {
"$ref": "#/$defs/DynamicRenameInput"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeDynamicRename",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
dynamic_rename_input(DynamicRenameInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1902 1903 1904 1905 1906 1907 1908 1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 | |
get_default_description()
Describes the dynamic rename rule.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1909 1910 1911 1912 1913 1914 1915 1916 1917 1918 1919 1920 1921 1922 1923 1924 1925 1926 1927 1928 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeEvaluateModel
pydantic-model
Bases: NodeSingleInput
Compute model-quality metrics by comparing actual and predicted columns.
Show JSON schema:
{
"$defs": {
"EvaluateModelSettings": {
"description": "Settings payload for the Evaluate Model node.\n\nDecoupled from any specific Train/Apply pair: takes a dataframe that\nalready contains both the actual target column and a prediction column\nand emits a long-form ``(metric, value)`` frame. Reusable on training,\ntest, or hold-out splits.\n\n``task_type=\"auto\"`` resolves the metric set from an upstream Train\nModel node when one is configured; otherwise defaults to ``regression``.",
"properties": {
"actual_column": {
"default": "",
"title": "Actual Column",
"type": "string"
},
"predicted_column": {
"default": "prediction",
"title": "Predicted Column",
"type": "string"
},
"task_type": {
"default": "auto",
"enum": [
"auto",
"regression",
"classification"
],
"title": "Task Type",
"type": "string"
},
"upstream_train_node_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Upstream Train Node Id"
}
},
"title": "EvaluateModelSettings",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Compute model-quality metrics by comparing actual and predicted columns.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"evaluate_input": {
"$ref": "#/$defs/EvaluateModelSettings"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeEvaluateModel",
"type": "object"
}
Config:
protected_namespaces:()
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
evaluate_input(EvaluateModelSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeExploreData
pydantic-model
Bases: NodeBase
Settings for a node that provides an interactive data exploration interface.
Show JSON schema:
{
"$defs": {
"DataModel": {
"properties": {
"data": {
"items": {
"additionalProperties": true,
"type": "object"
},
"title": "Data",
"type": "array"
},
"fields": {
"items": {
"$ref": "#/$defs/MutField"
},
"title": "Fields",
"type": "array"
}
},
"required": [
"data",
"fields"
],
"title": "DataModel",
"type": "object"
},
"GraphicWalkerInput": {
"properties": {
"dataModel": {
"$ref": "#/$defs/DataModel"
},
"is_initial": {
"default": true,
"title": "Is Initial",
"type": "boolean"
},
"specList": {
"anyOf": [
{
"items": {},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Speclist"
}
},
"title": "GraphicWalkerInput",
"type": "object"
},
"MutField": {
"properties": {
"fid": {
"title": "Fid",
"type": "string"
},
"key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Key"
},
"name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Name"
},
"basename": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Basename"
},
"disable": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Disable"
},
"semanticType": {
"title": "Semantictype",
"type": "string"
},
"analyticType": {
"enum": [
"measure",
"dimension"
],
"title": "Analytictype",
"type": "string"
},
"path": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Path"
},
"offset": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Offset"
}
},
"required": [
"fid",
"semanticType",
"analyticType"
],
"title": "MutField",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that provides an interactive data exploration interface.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"graphic_walker_input": {
"anyOf": [
{
"$ref": "#/$defs/GraphicWalkerInput"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeExploreData",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
graphic_walker_input(GraphicWalkerInput | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1863 1864 1865 1866 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeExternalSource
pydantic-model
Bases: NodeBase
Settings for a node that connects to a registered external data source.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SampleUsers": {
"description": "Settings for generating a sample dataset of users.",
"properties": {
"orientation": {
"default": "row",
"title": "Orientation",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
},
"SAMPLE_USERS": {
"title": "Sample Users",
"type": "boolean"
},
"class_name": {
"default": "sample_users",
"title": "Class Name",
"type": "string"
},
"size": {
"default": 100,
"title": "Size",
"type": "integer"
}
},
"required": [
"SAMPLE_USERS"
],
"title": "SampleUsers",
"type": "object"
}
},
"description": "Settings for a node that connects to a registered external data source.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"identifier": {
"title": "Identifier",
"type": "string"
},
"source_settings": {
"$ref": "#/$defs/SampleUsers"
}
},
"required": [
"flow_id",
"node_id",
"identifier",
"source_settings"
],
"title": "NodeExternalSource",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
identifier(str) -
source_settings(SampleUsers)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1150 1151 1152 1153 1154 1155 1156 1157 1158 | |
get_default_description()
Describes the external source.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1156 1157 1158 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeFilter
pydantic-model
Bases: NodeSingleInput
Settings for a node that filters rows based on a condition.
Show JSON schema:
{
"$defs": {
"BasicFilter": {
"description": "Defines a simple, single-condition filter (e.g., 'column' 'equals' 'value').\n\nAttributes:\n field: The column name to filter on.\n operator: The comparison operator (FilterOperator enum value or symbol).\n value: The value to compare against.\n value2: Second value for BETWEEN operator (optional).",
"properties": {
"field": {
"default": "",
"title": "Field",
"type": "string"
},
"operator": {
"anyOf": [
{
"$ref": "#/$defs/FilterOperator"
},
{
"type": "string"
}
],
"default": "equals",
"title": "Operator"
},
"value": {
"default": "",
"title": "Value",
"type": "string"
},
"value2": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Value2"
},
"filter_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Type"
},
"filter_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Value"
}
},
"title": "BasicFilter",
"type": "object"
},
"FilterInput": {
"description": "Defines the settings for a filter operation, supporting basic or advanced (expression-based) modes.\n\nAttributes:\n mode: The filter mode - \"basic\" or \"advanced\".\n basic_filter: The basic filter configuration (used when mode=\"basic\").\n advanced_filter: The advanced filter expression string (used when mode=\"advanced\").",
"properties": {
"mode": {
"default": "basic",
"enum": [
"basic",
"advanced"
],
"title": "Mode",
"type": "string"
},
"basic_filter": {
"anyOf": [
{
"$ref": "#/$defs/BasicFilter"
},
{
"type": "null"
}
],
"default": null
},
"advanced_filter": {
"default": "",
"expression": true,
"title": "Advanced Filter",
"type": "string"
},
"filter_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Type"
}
},
"title": "FilterInput",
"type": "object"
},
"FilterOperator": {
"description": "Supported filter comparison operators.",
"enum": [
"equals",
"not_equals",
"greater_than",
"greater_than_or_equals",
"less_than",
"less_than_or_equals",
"contains",
"not_contains",
"starts_with",
"ends_with",
"is_null",
"is_not_null",
"in",
"not_in",
"between"
],
"title": "FilterOperator",
"type": "string"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that filters rows based on a condition.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"filter_input": {
"$ref": "#/$defs/FilterInput"
},
"split_mode": {
"default": false,
"title": "Split Mode",
"type": "boolean"
}
},
"required": [
"flow_id",
"node_id",
"filter_input"
],
"title": "NodeFilter",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
filter_input(FilterInput) -
split_mode(bool)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 | |
output_names
property
Declared output handles, so the canvas can build both without the drawer.
The template's output stays 1 for backwards compatibility; the canvas
sizes the handle list off max(output, len(output_names)). Matches the
keys FlowDataEngine.filter_split produces at run time.
get_default_description()
Describes the filter condition.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeFlowInput
pydantic-model
Bases: NodeManualInput
Named source placeholder inside a subflow.
Extends NodeManualInput so the sample data reuses the manual-input settings
shape and editor. Standalone runs serve raw_data_format (empty frame when
blank); a parent run_flow node injects real data at execution time.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"RawData": {
"description": "Represents data in a raw, columnar format for manual input.",
"properties": {
"columns": {
"description": "Schema in column order. The i-th MinimalFieldInfo describes the values in data[i].",
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"title": "Columns",
"type": "array"
},
"data": {
"description": "Columnar layout: data[i] is the list of values for columns[i], in column order. len(data) must equal len(columns); each inner list has the same length (one entry per row). For two rows of {name, age}, emit [[\"Alice\", \"Bob\"], [30, 25]] \u2014 NOT [[\"Alice\", 30], [\"Bob\", 25]]. Reading rows back is `data[col_idx][row_idx]`.",
"items": {
"items": {},
"type": "array"
},
"title": "Data",
"type": "array"
}
},
"required": [
"columns",
"data"
],
"title": "RawData",
"type": "object"
}
},
"description": "Named source placeholder inside a subflow.\n\nExtends NodeManualInput so the sample data reuses the manual-input settings\nshape and editor. Standalone runs serve ``raw_data_format`` (empty frame when\nblank); a parent run_flow node injects real data at execution time.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"raw_data_format": {
"$ref": "#/$defs/RawData"
},
"input_name": {
"default": "input",
"title": "Input Name",
"type": "string"
}
},
"required": [
"flow_id",
"node_id",
"raw_data_format"
],
"title": "NodeFlowInput",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
raw_data_format(RawData) -
input_name(str)
Validators:
-
validate_node_reference→node_reference -
_coerce_none_raw_data_format -
_validate_input_name→input_name
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeFlowOutput
pydantic-model
Bases: NodeSingleInput
Named passthrough sink marking a subflow output; multiple allowed per flow.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Named passthrough sink marking a subflow output; multiple allowed per flow.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"output_name": {
"default": "output",
"title": "Output Name",
"type": "string"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeFlowOutput",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
output_name(str)
Validators:
-
validate_node_reference→node_reference -
_validate_output_name→output_name
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeFormula
pydantic-model
Bases: NodeSingleInput
Settings for a node that applies a formula to create/modify a column.
Show JSON schema:
{
"$defs": {
"DataType": {
"description": "Specific data types for fine-grained control.",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Categorical",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array"
],
"title": "DataType",
"type": "string"
},
"FieldInput": {
"description": "Represents a single field with its name and data type, typically for defining an output column.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"anyOf": [
{
"$ref": "#/$defs/DataType"
},
{
"const": "Auto",
"type": "string"
},
{
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "Auto",
"title": "Data Type"
}
},
"required": [
"name"
],
"title": "FieldInput",
"type": "object"
},
"FunctionInput": {
"description": "Defines a formula to be applied, including the output field information.",
"properties": {
"field": {
"$ref": "#/$defs/FieldInput"
},
"function": {
"expression": true,
"title": "Function",
"type": "string"
}
},
"required": [
"field",
"function"
],
"title": "FunctionInput",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that applies a formula to create/modify a column.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"function": {
"$ref": "#/$defs/FunctionInput",
"default": null
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeFormula",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
function(FunctionInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1334 1335 1336 1337 1338 1339 1340 1341 1342 1343 1344 1345 1346 1347 | |
get_default_description()
Describes the formula being applied.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1339 1340 1341 1342 1343 1344 1345 1346 1347 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeFuzzyMatch
pydantic-model
Bases: NodeJoin
Settings for a node that performs a fuzzy join based on string similarity.
Show JSON schema:
{
"$defs": {
"FuzzyMapping": {
"description": "Represents the configuration for a fuzzy string match between two columns.\n\nThis class defines all the necessary parameters to perform a fuzzy join,\nincluding the columns to match, the specific algorithm to use, and the\nsimilarity threshold required to consider two strings a match.\n\nIt generates a default name for the output score column if one is not\nprovided.\n\nAttributes:\n left_col (str): The name of the column in the left dataframe to join on.\n right_col (str): The name of the column in the right dataframe to join on.\n threshold_score (float): The similarity score threshold required for a\n match, typically on a scale of 0 to 100. Defaults to 80.0.\n fuzzy_type (FuzzyTypeLiteral): The string-matching algorithm to use.\n Defaults to \"levenshtein\".\n perc_unique (float): A parameter that may be used to assess column\n uniqueness before performing a costly fuzzy match. Defaults to 0.0.\n output_column_name (str | None): The name for the new column that will\n contain the calculated fuzzy match score. If None, a name is\n generated automatically in the format 'fuzzy_score_{left_col}_{right_col}'.\n valid (bool): A flag to indicate whether this mapping is active and should\n be used in a join operation. Defaults to True.\n reversed_threshold_score (float): A property that converts the 0-100\n threshold score into a 0.0-1.0 distance score, where 0.0 is a\n perfect match.",
"properties": {
"left_col": {
"title": "Left Col",
"type": "string"
},
"right_col": {
"title": "Right Col",
"type": "string"
},
"threshold_score": {
"default": 80.0,
"title": "Threshold Score",
"type": "number"
},
"fuzzy_type": {
"default": "levenshtein",
"enum": [
"levenshtein",
"jaro",
"jaro_winkler",
"hamming",
"damerau_levenshtein",
"indel"
],
"title": "Fuzzy Type",
"type": "string"
},
"perc_unique": {
"default": 0.0,
"title": "Perc Unique",
"type": "number"
},
"output_column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Column Name"
},
"valid": {
"default": true,
"title": "Valid",
"type": "boolean"
}
},
"required": [
"left_col",
"right_col"
],
"title": "FuzzyMapping",
"type": "object"
},
"FuzzyMatchInput": {
"description": "Data model for fuzzy matching join operations.",
"properties": {
"join_mapping": {
"items": {
"$ref": "#/$defs/FuzzyMapping"
},
"title": "Join Mapping",
"type": "array"
},
"left_select": {
"$ref": "#/$defs/JoinInputs"
},
"right_select": {
"$ref": "#/$defs/JoinInputs"
},
"how": {
"default": "inner",
"enum": [
"inner",
"left",
"right",
"full",
"semi",
"anti",
"cross",
"outer"
],
"title": "How",
"type": "string"
},
"aggregate_output": {
"default": false,
"title": "Aggregate Output",
"type": "boolean"
}
},
"required": [
"join_mapping",
"left_select",
"right_select"
],
"title": "FuzzyMatchInput",
"type": "object"
},
"JoinInputs": {
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Settings for a node that performs a fuzzy join based on string similarity.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"auto_generate_selection": {
"default": true,
"title": "Auto Generate Selection",
"type": "boolean"
},
"verify_integrity": {
"default": true,
"title": "Verify Integrity",
"type": "boolean"
},
"join_input": {
"$ref": "#/$defs/FuzzyMatchInput"
},
"auto_keep_all": {
"default": true,
"title": "Auto Keep All",
"type": "boolean"
},
"auto_keep_right": {
"default": true,
"title": "Auto Keep Right",
"type": "boolean"
},
"auto_keep_left": {
"default": true,
"title": "Auto Keep Left",
"type": "boolean"
}
},
"required": [
"flow_id",
"node_id",
"join_input"
],
"title": "NodeFuzzyMatch",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
auto_generate_selection(bool) -
verify_integrity(bool) -
auto_keep_all(bool) -
auto_keep_right(bool) -
auto_keep_left(bool) -
join_input(FuzzyMatchInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 | |
get_default_description()
Describes the fuzzy match join.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
810 811 812 813 814 815 816 817 818 819 820 821 822 823 | |
to_yaml_dict()
Converts the fuzzy match node settings to a dictionary for YAML serialization.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeGoogleAnalyticsReader
pydantic-model
Bases: NodeBase
Settings for a node that reads from a Google Analytics 4 property.
Show JSON schema:
{
"$defs": {
"GoogleAnalyticsFilter": {
"description": "A single filter applied to a GA4 dimension or metric.\n\n``field`` must match one of the selected dimensions or metrics; the worker\nauto-routes the filter into either the request's ``dimension_filter`` (for\nstring-typed dimensions) or ``metric_filter`` (for numeric-typed metrics).\n\nSupported operators \u2014 strings (dimensions):\n - ``equals``, ``not_equals``\n - ``contains``, ``begins_with``, ``ends_with``\n - ``regex`` (full regex match)\n - ``in_list``, ``not_in_list`` (comma-separated ``value``)\n\nSupported operators \u2014 numeric (metrics):\n - ``equals``, ``not_equals``\n - ``less_than``, ``less_equal``, ``greater_than``, ``greater_equal``\n - ``between`` (comma-separated ``\"low,high\"``)\n\nMultiple filters on the same kind are AND-combined. String matching is\ncase-insensitive by default (``case_sensitive=False`` below).",
"properties": {
"field": {
"title": "Field",
"type": "string"
},
"operator": {
"title": "Operator",
"type": "string"
},
"value": {
"default": "",
"title": "Value",
"type": "string"
},
"case_sensitive": {
"default": false,
"title": "Case Sensitive",
"type": "boolean"
}
},
"required": [
"field",
"operator"
],
"title": "GoogleAnalyticsFilter",
"type": "object"
},
"GoogleAnalyticsOrderBy": {
"description": "A single sort entry applied to the GA4 report.\n\n``field`` must match one of the selected dimensions or metrics; the worker\nroutes it into a ``DimensionOrderBy`` or ``MetricOrderBy`` accordingly.\n``descending=True`` produces a descending sort. Sort entries are applied in\nlist order.",
"properties": {
"field": {
"title": "Field",
"type": "string"
},
"descending": {
"default": false,
"title": "Descending",
"type": "boolean"
}
},
"required": [
"field"
],
"title": "GoogleAnalyticsOrderBy",
"type": "object"
},
"GoogleAnalyticsSettings": {
"description": "UI settings for a Google Analytics 4 reader node.\n\nCredentials are NOT stored inline: ``ga_connection_name`` is a reference to\na Google Analytics connection managed under ``/ga_connections`` (whose\nservice-account JSON is encrypted at rest).",
"properties": {
"ga_connection_name": {
"title": "Ga Connection Name",
"type": "string"
},
"property_id": {
"title": "Property Id",
"type": "string"
},
"start_date": {
"default": "7daysAgo",
"title": "Start Date",
"type": "string"
},
"end_date": {
"default": "yesterday",
"title": "End Date",
"type": "string"
},
"metrics": {
"items": {
"type": "string"
},
"title": "Metrics",
"type": "array"
},
"dimensions": {
"items": {
"type": "string"
},
"title": "Dimensions",
"type": "array"
},
"limit": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Limit"
},
"filters": {
"items": {
"$ref": "#/$defs/GoogleAnalyticsFilter"
},
"title": "Filters",
"type": "array"
},
"order_bys": {
"items": {
"$ref": "#/$defs/GoogleAnalyticsOrderBy"
},
"title": "Order Bys",
"type": "array"
}
},
"required": [
"ga_connection_name",
"property_id"
],
"title": "GoogleAnalyticsSettings",
"type": "object"
},
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that reads from a Google Analytics 4 property.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"google_analytics_settings": {
"$ref": "#/$defs/GoogleAnalyticsSettings"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
}
},
"required": [
"flow_id",
"node_id",
"google_analytics_settings"
],
"title": "NodeGoogleAnalyticsReader",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
google_analytics_settings(GoogleAnalyticsSettings) -
fields(list[MinimalFieldInfo] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 | |
get_default_description()
Describes the GA4 query.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeGraphSolver
pydantic-model
Bases: NodeSingleInput
Settings for a node that solves graph-based problems (e.g., connected components).
Show JSON schema:
{
"$defs": {
"GraphSolverInput": {
"description": "Defines settings for a graph-solving operation (e.g., finding connected components).",
"properties": {
"col_from": {
"title": "Col From",
"type": "string"
},
"col_to": {
"title": "Col To",
"type": "string"
},
"output_column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "graph_group",
"title": "Output Column Name"
}
},
"required": [
"col_from",
"col_to"
],
"title": "GraphSolverInput",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that solves graph-based problems (e.g., connected components).",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"graph_solver_input": {
"$ref": "#/$defs/GraphSolverInput"
}
},
"required": [
"flow_id",
"node_id",
"graph_solver_input"
],
"title": "NodeGraphSolver",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
graph_solver_input(GraphSolverInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1869 1870 1871 1872 1873 1874 1875 1876 1877 | |
get_default_description()
Describes the graph solver operation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1874 1875 1876 1877 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeGroupBy
pydantic-model
Bases: NodeSingleInput
Settings for a node that performs a group-by and aggregation operation.
Show JSON schema:
{
"$defs": {
"AggColl": {
"description": "A data class that represents a single aggregation operation for a group by operation.\n\nAttributes\n----------\nold_name : str\n The name of the column in the original DataFrame to be aggregated.\n\nagg : str\n The aggregation function to use. This can be a string representing a built-in function or a custom function.\n\nnew_name : Optional[str]\n The name of the resulting aggregated column in the output DataFrame. If not provided, it will default to the\n old_name appended with the aggregation function.\n\noutput_type : Optional[str]\n The type of the output values of the aggregation. If not provided, it is inferred from the aggregation function\n using the `get_func_type_mapping` function.\n\nExample\n--------\nagg_col = AggColl(\n old_name='col1',\n agg='sum',\n new_name='sum_col1',\n output_type='float'\n)",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"agg": {
"title": "Agg",
"type": "string"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"output_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Type"
}
},
"required": [
"old_name",
"agg"
],
"title": "AggColl",
"type": "object"
},
"GroupByInput": {
"description": "A data class that represents the input for a group by operation.\n\nAttributes\n----------\nagg_cols : List[AggColl]\n A list of `AggColl` objects that specify the aggregation operations to perform on the DataFrame columns\n after grouping. Each `AggColl` object should specify the column to be aggregated and the aggregation\n function to use.\n\nExample\n--------\ngroup_by_input = GroupByInput(\n agg_cols=[AggColl(old_name='ix', agg='groupby'), AggColl(old_name='groups', agg='groupby'),\n AggColl(old_name='col1', agg='sum'), AggColl(old_name='col2', agg='mean')]\n)",
"properties": {
"agg_cols": {
"items": {
"$ref": "#/$defs/AggColl"
},
"title": "Agg Cols",
"type": "array"
}
},
"required": [
"agg_cols"
],
"title": "GroupByInput",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that performs a group-by and aggregation operation.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"groupby_input": {
"$ref": "#/$defs/GroupByInput",
"default": null
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeGroupBy",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
groupby_input(GroupByInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 | |
get_default_description()
Describes the group-by columns and aggregations.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 1392 1393 1394 1395 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeInputConnection
pydantic-model
Bases: BaseModel
Represents the input side of a connection between two nodes.
Show JSON schema:
{
"description": "Represents the input side of a connection between two nodes.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"connection_class": {
"enum": [
"input-0",
"input-1",
"input-2",
"input-3",
"input-4",
"input-5",
"input-6",
"input-7",
"input-8",
"input-9"
],
"title": "Connection Class",
"type": "string"
}
},
"required": [
"node_id",
"connection_class"
],
"title": "NodeInputConnection",
"type": "object"
}
Fields:
-
node_id(int) -
connection_class(InputConnectionClass)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 | |
get_node_input_connection_type()
Determines the semantic type of the input (e.g., for a join).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 | |
NodeJoin
pydantic-model
Bases: NodeMultiInput
Settings for a node that performs a standard SQL-style join.
Show JSON schema:
{
"$defs": {
"JoinInput": {
"description": "Data model for standard SQL-style join operations.",
"properties": {
"join_mapping": {
"items": {
"$ref": "#/$defs/JoinMap"
},
"title": "Join Mapping",
"type": "array"
},
"left_select": {
"$ref": "#/$defs/JoinInputs"
},
"right_select": {
"$ref": "#/$defs/JoinInputs"
},
"how": {
"default": "inner",
"enum": [
"inner",
"left",
"right",
"full",
"semi",
"anti",
"outer"
],
"title": "How",
"type": "string"
}
},
"required": [
"join_mapping",
"left_select",
"right_select"
],
"title": "JoinInput",
"type": "object"
},
"JoinInputs": {
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
},
"JoinMap": {
"description": "Defines a single mapping between a left and right column for a join key.",
"properties": {
"left_col": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Col"
},
"right_col": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Col"
}
},
"title": "JoinMap",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Settings for a node that performs a standard SQL-style join.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"auto_generate_selection": {
"default": true,
"title": "Auto Generate Selection",
"type": "boolean"
},
"verify_integrity": {
"default": true,
"title": "Verify Integrity",
"type": "boolean"
},
"join_input": {
"$ref": "#/$defs/JoinInput"
},
"auto_keep_all": {
"default": true,
"title": "Auto Keep All",
"type": "boolean"
},
"auto_keep_right": {
"default": true,
"title": "Auto Keep Right",
"type": "boolean"
},
"auto_keep_left": {
"default": true,
"title": "Auto Keep Left",
"type": "boolean"
}
},
"required": [
"flow_id",
"node_id",
"join_input"
],
"title": "NodeJoin",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
auto_generate_selection(bool) -
verify_integrity(bool) -
join_input(JoinInput) -
auto_keep_all(bool) -
auto_keep_right(bool) -
auto_keep_left(bool)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 | |
get_default_description()
Describes the join type and key columns.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
720 721 722 723 724 725 726 727 728 729 730 731 732 733 | |
to_yaml_dict()
Converts the join node settings to a dictionary for YAML serialization.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeKafkaSource
pydantic-model
Bases: NodeBase
Settings for a node that reads from a Kafka or Redpanda topic.
Show JSON schema:
{
"$defs": {
"KafkaSourceSettings": {
"description": "Configuration for reading from a Kafka/Redpanda topic.",
"properties": {
"kafka_connection_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Kafka Connection Id"
},
"kafka_connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Kafka Connection Name"
},
"topic_name": {
"default": "",
"title": "Topic Name",
"type": "string"
},
"value_format": {
"const": "json",
"default": "json",
"title": "Value Format",
"type": "string"
},
"sync_name": {
"default": "",
"title": "Sync Name",
"type": "string"
},
"start_offset": {
"default": "latest",
"enum": [
"earliest",
"latest"
],
"title": "Start Offset",
"type": "string"
},
"max_messages": {
"default": 100000,
"title": "Max Messages",
"type": "integer"
},
"poll_timeout_seconds": {
"default": 30.0,
"title": "Poll Timeout Seconds",
"type": "number"
}
},
"title": "KafkaSourceSettings",
"type": "object"
},
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that reads from a Kafka or Redpanda topic.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"kafka_settings": {
"$ref": "#/$defs/KafkaSourceSettings",
"default": {
"kafka_connection_id": null,
"kafka_connection_name": null,
"topic_name": "",
"value_format": "json",
"sync_name": "",
"start_offset": "latest",
"max_messages": 100000,
"poll_timeout_seconds": 30.0
}
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeKafkaSource",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
kafka_settings(KafkaSourceSettings) -
fields(list[MinimalFieldInfo] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1808 1809 1810 1811 1812 1813 1814 1815 1816 1817 1818 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeManualInput
pydantic-model
Bases: NodeBase
Settings for a node that allows direct data entry in the UI.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"RawData": {
"description": "Represents data in a raw, columnar format for manual input.",
"properties": {
"columns": {
"description": "Schema in column order. The i-th MinimalFieldInfo describes the values in data[i].",
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"title": "Columns",
"type": "array"
},
"data": {
"description": "Columnar layout: data[i] is the list of values for columns[i], in column order. len(data) must equal len(columns); each inner list has the same length (one entry per row). For two rows of {name, age}, emit [[\"Alice\", \"Bob\"], [30, 25]] \u2014 NOT [[\"Alice\", 30], [\"Bob\", 25]]. Reading rows back is `data[col_idx][row_idx]`.",
"items": {
"items": {},
"type": "array"
},
"title": "Data",
"type": "array"
}
},
"required": [
"columns",
"data"
],
"title": "RawData",
"type": "object"
}
},
"description": "Settings for a node that allows direct data entry in the UI.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"raw_data_format": {
"$ref": "#/$defs/RawData"
}
},
"required": [
"flow_id",
"node_id",
"raw_data_format"
],
"title": "NodeManualInput",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
raw_data_format(RawData)
Validators:
-
validate_node_reference→node_reference -
_coerce_none_raw_data_format
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 | |
get_default_description()
Describes the manual input columns.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
930 931 932 933 934 935 936 937 938 939 940 941 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeMultiInput
pydantic-model
Bases: NodeBase
A base model for any node that takes multiple data inputs.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "A base model for any node that takes multiple data inputs.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeMultiInput",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
481 482 483 484 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeOutput
pydantic-model
Bases: NodeSingleInput
Settings for a node that writes its input to a file.
Show JSON schema:
{
"$defs": {
"OutputAvroTable": {
"description": "Defines settings for writing an Avro file.",
"properties": {
"file_type": {
"const": "avro",
"default": "avro",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"snappy",
"deflate"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputAvroTable",
"type": "object"
},
"OutputCsvTable": {
"description": "Defines settings for writing a CSV file.",
"properties": {
"file_type": {
"const": "csv",
"default": "csv",
"title": "File Type",
"type": "string"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
}
},
"title": "OutputCsvTable",
"type": "object"
},
"OutputExcelTable": {
"description": "Defines settings for writing an Excel file.",
"properties": {
"file_type": {
"const": "excel",
"default": "excel",
"title": "File Type",
"type": "string"
},
"sheet_name": {
"default": "Sheet1",
"title": "Sheet Name",
"type": "string"
}
},
"title": "OutputExcelTable",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"OutputIpcTable": {
"description": "Defines settings for writing an Arrow IPC/Feather file.",
"properties": {
"file_type": {
"const": "ipc",
"default": "ipc",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"lz4",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputIpcTable",
"type": "object"
},
"OutputNdjsonTable": {
"description": "Defines settings for writing a newline-delimited JSON file.",
"properties": {
"file_type": {
"const": "ndjson",
"default": "ndjson",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"gzip",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputNdjsonTable",
"type": "object"
},
"OutputParquetTable": {
"description": "Defines settings for writing a Parquet file.",
"properties": {
"file_type": {
"const": "parquet",
"default": "parquet",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "zstd",
"enum": [
"lz4",
"uncompressed",
"snappy",
"gzip",
"brotli",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputParquetTable",
"type": "object"
},
"OutputSettings": {
"description": "Defines the complete settings for an output node.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"directory": {
"title": "Directory",
"type": "string"
},
"file_type": {
"title": "File Type",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Fields"
},
"write_mode": {
"default": "overwrite",
"title": "Write Mode",
"type": "string"
},
"table_settings": {
"discriminator": {
"mapping": {
"avro": "#/$defs/OutputAvroTable",
"csv": "#/$defs/OutputCsvTable",
"excel": "#/$defs/OutputExcelTable",
"ipc": "#/$defs/OutputIpcTable",
"ndjson": "#/$defs/OutputNdjsonTable",
"parquet": "#/$defs/OutputParquetTable"
},
"propertyName": "file_type"
},
"oneOf": [
{
"$ref": "#/$defs/OutputCsvTable"
},
{
"$ref": "#/$defs/OutputParquetTable"
},
{
"$ref": "#/$defs/OutputExcelTable"
},
{
"$ref": "#/$defs/OutputIpcTable"
},
{
"$ref": "#/$defs/OutputNdjsonTable"
},
{
"$ref": "#/$defs/OutputAvroTable"
}
],
"title": "Table Settings"
},
"abs_file_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Abs File Path"
}
},
"required": [
"name",
"directory",
"file_type",
"table_settings"
],
"title": "OutputSettings",
"type": "object"
}
},
"description": "Settings for a node that writes its input to a file.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"output_settings": {
"$ref": "#/$defs/OutputSettings"
}
},
"required": [
"flow_id",
"node_id",
"output_settings"
],
"title": "NodeOutput",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
output_settings(OutputSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 | |
get_default_description()
Describes the output file target.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1476 1477 1478 1479 | |
to_yaml_dict()
Converts the output node settings to a dictionary for YAML serialization.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeOutputConnection
pydantic-model
Bases: BaseModel
Represents the output side of a connection between two nodes.
Show JSON schema:
{
"description": "Represents the output side of a connection between two nodes.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"connection_class": {
"enum": [
"output-0",
"output-1",
"output-2",
"output-3",
"output-4",
"output-5",
"output-6",
"output-7",
"output-8",
"output-9"
],
"title": "Connection Class",
"type": "string"
}
},
"required": [
"node_id",
"connection_class"
],
"title": "NodeOutputConnection",
"type": "object"
}
Fields:
-
node_id(int) -
connection_class(OutputConnectionClass)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1821 1822 1823 1824 1825 | |
NodePivot
pydantic-model
Bases: NodeSingleInput
Settings for a node that pivots data from a long to a wide format.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"PivotInput": {
"description": "Defines the settings for a pivot (long-to-wide) operation.",
"properties": {
"index_columns": {
"items": {
"type": "string"
},
"title": "Index Columns",
"type": "array"
},
"pivot_column": {
"title": "Pivot Column",
"type": "string"
},
"value_col": {
"title": "Value Col",
"type": "string"
},
"aggregations": {
"items": {
"type": "string"
},
"title": "Aggregations",
"type": "array"
}
},
"required": [
"index_columns",
"pivot_column",
"value_col",
"aggregations"
],
"title": "PivotInput",
"type": "object"
}
},
"description": "Settings for a node that pivots data from a long to a wide format.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"pivot_input": {
"$ref": "#/$defs/PivotInput",
"default": null
},
"output_fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Fields"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodePivot",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
pivot_input(PivotInput) -
output_fields(list[MinimalFieldInfo] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 | |
get_default_description()
Describes the pivot operation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1430 1431 1432 1433 1434 1435 1436 1437 1438 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodePolarsCode
pydantic-model
Bases: NodeMultiInput
Settings for a node that executes arbitrary user-provided Polars code.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"PolarsCodeInput": {
"description": "A simple container for a string of user-provided Polars code to be executed.",
"properties": {
"polars_code": {
"title": "Polars Code",
"type": "string"
}
},
"required": [
"polars_code"
],
"title": "PolarsCodeInput",
"type": "object"
}
},
"description": "Settings for a node that executes arbitrary user-provided Polars code.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"polars_code_input": {
"$ref": "#/$defs/PolarsCodeInput"
}
},
"required": [
"flow_id",
"node_id",
"polars_code_input"
],
"title": "NodePolarsCode",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
polars_code_input(PolarsCodeInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1931 1932 1933 1934 1935 1936 1937 1938 1939 1940 1941 1942 | |
get_default_description()
Describes the Polars code snippet.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1936 1937 1938 1939 1940 1941 1942 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodePromise
pydantic-model
Bases: NodeBase
A placeholder node for an operation that has not yet been configured.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "A placeholder node for an operation that has not yet been configured.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"default": false,
"title": "Is Setup",
"type": "boolean"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"node_type": {
"title": "Node Type",
"type": "string"
}
},
"required": [
"flow_id",
"node_id",
"node_type"
],
"title": "NodePromise",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
is_setup(bool) -
node_type(str)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1398 1399 1400 1401 1402 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodePythonScript
pydantic-model
Bases: NodeMultiInput
Node that executes Python code on a kernel container.
Show JSON schema:
{
"$defs": {
"NotebookCell": {
"description": "A single cell in the notebook editor.\n\nNote: Cell output (stdout, display_outputs, errors) is handled entirely\non the frontend and is not persisted. Only id and code are stored.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"code": {
"default": "",
"title": "Code",
"type": "string"
}
},
"required": [
"id"
],
"title": "NotebookCell",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"PythonScriptInput": {
"description": "Settings for Python code execution on a kernel.",
"properties": {
"code": {
"default": "",
"title": "Code",
"type": "string"
},
"kernel_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Kernel Id"
},
"cells": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/NotebookCell"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Cells"
}
},
"title": "PythonScriptInput",
"type": "object"
}
},
"description": "Node that executes Python code on a kernel container.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"python_script_input": {
"$ref": "#/$defs/PythonScriptInput",
"default": {
"code": "",
"kernel_id": null,
"cells": null
}
},
"output_names": {
"items": {
"type": "string"
},
"title": "Output Names",
"type": "array"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodePythonScript",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
python_script_input(PythonScriptInput) -
output_names(list[str])
Validators:
-
validate_node_reference→node_reference -
validate_output_names→output_names
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeRandomSplit
pydantic-model
Bases: NodeSingleInput
Settings for a node that randomly partitions rows into N labeled outputs.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"RandomSplitGroup": {
"description": "A single output partition in a random split.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"percentage": {
"title": "Percentage",
"type": "number"
}
},
"required": [
"name",
"percentage"
],
"title": "RandomSplitGroup",
"type": "object"
}
},
"description": "Settings for a node that randomly partitions rows into N labeled outputs.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"splits": {
"items": {
"$ref": "#/$defs/RandomSplitGroup"
},
"title": "Splits",
"type": "array"
},
"seed": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Seed"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeRandomSplit",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
splits(list[RandomSplitGroup]) -
seed(int | None)
Validators:
-
validate_node_reference→node_reference -
_validate_splits
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeRead
pydantic-model
Bases: NodeBase
Settings for a node that reads data from a file.
Show JSON schema:
{
"$defs": {
"InputAvroTable": {
"description": "Defines settings for reading an Avro file.",
"properties": {
"file_type": {
"const": "avro",
"default": "avro",
"title": "File Type",
"type": "string"
}
},
"title": "InputAvroTable",
"type": "object"
},
"InputCsvTable": {
"description": "Defines settings for reading a CSV file.",
"properties": {
"file_type": {
"const": "csv",
"default": "csv",
"title": "File Type",
"type": "string"
},
"reference": {
"default": "",
"title": "Reference",
"type": "string"
},
"starting_from_line": {
"default": 0,
"title": "Starting From Line",
"type": "integer"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
},
"parquet_ref": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Parquet Ref"
},
"row_delimiter": {
"default": "\n",
"title": "Row Delimiter",
"type": "string"
},
"quote_char": {
"default": "\"",
"title": "Quote Char",
"type": "string"
},
"infer_schema_length": {
"default": 10000,
"title": "Infer Schema Length",
"type": "integer"
},
"infer_schema": {
"default": true,
"title": "Infer Schema",
"type": "boolean"
},
"truncate_ragged_lines": {
"default": false,
"title": "Truncate Ragged Lines",
"type": "boolean"
},
"ignore_errors": {
"default": false,
"title": "Ignore Errors",
"type": "boolean"
}
},
"title": "InputCsvTable",
"type": "object"
},
"InputExcelTable": {
"description": "Defines settings for reading an Excel file.",
"properties": {
"file_type": {
"const": "excel",
"default": "excel",
"title": "File Type",
"type": "string"
},
"sheet_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Sheet Name"
},
"start_row": {
"default": 0,
"title": "Start Row",
"type": "integer"
},
"start_column": {
"default": 0,
"title": "Start Column",
"type": "integer"
},
"end_row": {
"default": 0,
"title": "End Row",
"type": "integer"
},
"end_column": {
"default": 0,
"title": "End Column",
"type": "integer"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"type_inference": {
"default": false,
"title": "Type Inference",
"type": "boolean"
}
},
"title": "InputExcelTable",
"type": "object"
},
"InputIpcTable": {
"description": "Defines settings for reading an Arrow IPC/Feather file.",
"properties": {
"file_type": {
"const": "ipc",
"default": "ipc",
"title": "File Type",
"type": "string"
}
},
"title": "InputIpcTable",
"type": "object"
},
"InputJsonTable": {
"description": "Defines settings for reading a JSON file.",
"properties": {
"file_type": {
"const": "json",
"default": "json",
"title": "File Type",
"type": "string"
},
"reference": {
"default": "",
"title": "Reference",
"type": "string"
},
"starting_from_line": {
"default": 0,
"title": "Starting From Line",
"type": "integer"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
},
"parquet_ref": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Parquet Ref"
},
"row_delimiter": {
"default": "\n",
"title": "Row Delimiter",
"type": "string"
},
"quote_char": {
"default": "\"",
"title": "Quote Char",
"type": "string"
},
"infer_schema_length": {
"default": 10000,
"title": "Infer Schema Length",
"type": "integer"
},
"infer_schema": {
"default": true,
"title": "Infer Schema",
"type": "boolean"
},
"truncate_ragged_lines": {
"default": false,
"title": "Truncate Ragged Lines",
"type": "boolean"
},
"ignore_errors": {
"default": false,
"title": "Ignore Errors",
"type": "boolean"
}
},
"title": "InputJsonTable",
"type": "object"
},
"InputNdjsonTable": {
"description": "Defines settings for reading a newline-delimited JSON file.",
"properties": {
"file_type": {
"const": "ndjson",
"default": "ndjson",
"title": "File Type",
"type": "string"
}
},
"title": "InputNdjsonTable",
"type": "object"
},
"InputParquetTable": {
"description": "Defines settings for reading a Parquet file.",
"properties": {
"file_type": {
"const": "parquet",
"default": "parquet",
"title": "File Type",
"type": "string"
}
},
"title": "InputParquetTable",
"type": "object"
},
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"ReceivedTable": {
"description": "Model for defining a table received from an external source.",
"properties": {
"id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Id"
},
"name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Name"
},
"path": {
"title": "Path",
"type": "string"
},
"directory": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Directory"
},
"analysis_file_available": {
"default": false,
"title": "Analysis File Available",
"type": "boolean"
},
"status": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Status"
},
"fields": {
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"title": "Fields",
"type": "array"
},
"abs_file_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Abs File Path"
},
"file_type": {
"enum": [
"csv",
"json",
"parquet",
"excel",
"ipc",
"ndjson",
"avro"
],
"title": "File Type",
"type": "string"
},
"table_settings": {
"discriminator": {
"mapping": {
"avro": "#/$defs/InputAvroTable",
"csv": "#/$defs/InputCsvTable",
"excel": "#/$defs/InputExcelTable",
"ipc": "#/$defs/InputIpcTable",
"json": "#/$defs/InputJsonTable",
"ndjson": "#/$defs/InputNdjsonTable",
"parquet": "#/$defs/InputParquetTable"
},
"propertyName": "file_type"
},
"oneOf": [
{
"$ref": "#/$defs/InputCsvTable"
},
{
"$ref": "#/$defs/InputJsonTable"
},
{
"$ref": "#/$defs/InputParquetTable"
},
{
"$ref": "#/$defs/InputExcelTable"
},
{
"$ref": "#/$defs/InputIpcTable"
},
{
"$ref": "#/$defs/InputNdjsonTable"
},
{
"$ref": "#/$defs/InputAvroTable"
}
],
"title": "Table Settings"
}
},
"required": [
"path",
"file_type",
"table_settings"
],
"title": "ReceivedTable",
"type": "object"
}
},
"description": "Settings for a node that reads data from a file.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"received_file": {
"$ref": "#/$defs/ReceivedTable"
}
},
"required": [
"flow_id",
"node_id",
"received_file"
],
"title": "NodeRead",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
received_file(ReceivedTable)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
944 945 946 947 948 949 950 951 952 953 | |
get_default_description()
Describes the file being read.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
949 950 951 952 953 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeRecordCount
pydantic-model
Bases: NodeSingleInput
Settings for a node that counts the number of records.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that counts the number of records.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeRecordCount",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1896 1897 1898 1899 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeRecordId
pydantic-model
Bases: NodeSingleInput
Settings for a node that adds a unique record ID column.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"RecordIdInput": {
"description": "Defines settings for adding a record ID (row number) column to the data.",
"properties": {
"output_column_name": {
"default": "record_id",
"title": "Output Column Name",
"type": "string"
},
"offset": {
"default": 1,
"title": "Offset",
"type": "integer"
},
"group_by": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Group By"
},
"group_by_columns": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Group By Columns"
}
},
"title": "RecordIdInput",
"type": "object"
}
},
"description": "Settings for a node that adds a unique record ID column.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"record_id_input": {
"$ref": "#/$defs/RecordIdInput"
}
},
"required": [
"flow_id",
"node_id",
"record_id_input"
],
"title": "NodeRecordId",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
record_id_input(RecordIdInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
695 696 697 698 699 700 701 702 703 704 705 706 707 | |
get_default_description()
Describes the record ID column being added.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
700 701 702 703 704 705 706 707 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeRestApiReader
pydantic-model
Bases: NodeBase
Settings for a node that reads from a REST API.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"RestApiAuthSettings": {
"description": "Authentication settings for a REST API reader node.\n\nThe credential (API key / bearer token / basic password, per ``auth_type``)\nis NOT stored inline. ``secret_name`` references a secret in the user's\nsecret store \u2014 created once via the Secrets manager and reusable across\nnodes \u2014 mirroring how the database reader references a stored password. The\n``.flowfile`` persists only the reference name, never the credential itself.\n\n``secret`` is an optional inline plaintext for programmatic use\n(``flowfile_frame.read_api``); it is encrypted with the master key and\ncleared, never persisted.",
"properties": {
"auth_type": {
"default": "none",
"enum": [
"none",
"api_key",
"bearer",
"basic"
],
"title": "Auth Type",
"type": "string"
},
"api_key_name": {
"default": "X-API-Key",
"title": "Api Key Name",
"type": "string"
},
"api_key_location": {
"default": "header",
"enum": [
"header",
"query"
],
"title": "Api Key Location",
"type": "string"
},
"basic_username": {
"default": "",
"title": "Basic Username",
"type": "string"
},
"secret_name": {
"default": "",
"title": "Secret Name",
"type": "string"
},
"secret": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Secret"
}
},
"title": "RestApiAuthSettings",
"type": "object"
},
"RestApiPaginationSettings": {
"description": "Pagination strategy and parameters for a REST API reader node.",
"properties": {
"pagination_type": {
"default": "none",
"enum": [
"none",
"offset",
"page",
"cursor"
],
"title": "Pagination Type",
"type": "string"
},
"offset_param": {
"default": "offset",
"title": "Offset Param",
"type": "string"
},
"limit_param": {
"default": "limit",
"title": "Limit Param",
"type": "string"
},
"page_size": {
"default": 100,
"title": "Page Size",
"type": "integer"
},
"page_param": {
"default": "page",
"title": "Page Param",
"type": "string"
},
"start_page": {
"default": 1,
"title": "Start Page",
"type": "integer"
},
"cursor_param": {
"default": "cursor",
"title": "Cursor Param",
"type": "string"
},
"cursor_location": {
"default": "body",
"enum": [
"body",
"header"
],
"title": "Cursor Location",
"type": "string"
},
"cursor_response_path": {
"default": "",
"title": "Cursor Response Path",
"type": "string"
},
"initial_cursor": {
"default": "",
"title": "Initial Cursor",
"type": "string"
},
"max_pages": {
"default": 1000,
"title": "Max Pages",
"type": "integer"
},
"max_records": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Records"
},
"page_delay_seconds": {
"default": 0.0,
"title": "Page Delay Seconds",
"type": "number"
}
},
"title": "RestApiPaginationSettings",
"type": "object"
},
"RestApiSettings": {
"description": "UI settings for a REST API reader node.\n\nSecrets are stored inline but encrypted (see ``RestApiAuthSettings``). JSON\nis the only supported response format; ``record_path`` is a dot-path that\nlocates the record array within the response body (empty = top-level).",
"properties": {
"url": {
"default": "",
"title": "Url",
"type": "string"
},
"method": {
"default": "GET",
"enum": [
"GET",
"POST"
],
"title": "Method",
"type": "string"
},
"headers": {
"additionalProperties": {
"type": "string"
},
"title": "Headers",
"type": "object"
},
"query_params": {
"additionalProperties": {
"type": "string"
},
"title": "Query Params",
"type": "object"
},
"json_body": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Json Body"
},
"auth": {
"$ref": "#/$defs/RestApiAuthSettings"
},
"pagination": {
"$ref": "#/$defs/RestApiPaginationSettings"
},
"record_path": {
"default": "",
"title": "Record Path",
"type": "string"
},
"timeout_seconds": {
"default": 30.0,
"title": "Timeout Seconds",
"type": "number"
},
"max_retries": {
"default": 3,
"title": "Max Retries",
"type": "integer"
}
},
"title": "RestApiSettings",
"type": "object"
}
},
"description": "Settings for a node that reads from a REST API.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"rest_api_settings": {
"$ref": "#/$defs/RestApiSettings"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
}
},
"required": [
"flow_id",
"node_id",
"rest_api_settings"
],
"title": "NodeRestApiReader",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
rest_api_settings(RestApiSettings) -
fields(list[MinimalFieldInfo] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 | |
get_default_description()
Describes the REST API request.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1323 1324 1325 1326 1327 1328 1329 1330 1331 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeRunFlow
pydantic-model
Bases: NodeBase
Settings for a node that executes a catalog-registered flow as a subflow.
input_slots/output_slots persist the last-synced subflow interface
(flow_input/flow_output names in interface order); connections are keyed to
handles positionally against them (handle input-{i+1} <-> input_slots[i];
input-0 is the reserved parameter-data handle).
Show JSON schema:
{
"$defs": {
"FlowParameter": {
"description": "A single flow-level parameter that can be referenced via ${name} syntax.\n\n``default_value`` stays a string for file-format stability; ``typed_default``\nyields the coerced Python value used for whole-field ``${name}`` injection.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"default_value": {
"default": "",
"title": "Default Value",
"type": "string"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"type": {
"default": "string",
"enum": [
"string",
"integer",
"float",
"boolean",
"enum"
],
"title": "Type",
"type": "string"
},
"enum_values": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Enum Values"
}
},
"required": [
"name"
],
"title": "FlowParameter",
"type": "object"
},
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"RunFlowParameterBinding": {
"description": "How one subflow parameter gets its value for a run_flow execution.",
"properties": {
"parameter_name": {
"title": "Parameter Name",
"type": "string"
},
"source": {
"default": "default",
"enum": [
"default",
"constant",
"column"
],
"title": "Source",
"type": "string"
},
"constant_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Constant Value"
},
"column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Column Name"
}
},
"required": [
"parameter_name"
],
"title": "RunFlowParameterBinding",
"type": "object"
},
"SubflowReference": {
"description": "Reference to a catalog-registered flow.\n\n``registration_id`` is the primary reference; ``flow_uuid`` is stamped\nserver-side and used to repair a dangling id; ``flow_path`` is display-only.",
"properties": {
"registration_id": {
"title": "Registration Id",
"type": "integer"
},
"flow_uuid": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Flow Uuid"
},
"flow_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Flow Path"
}
},
"required": [
"registration_id"
],
"title": "SubflowReference",
"type": "object"
}
},
"description": "Settings for a node that executes a catalog-registered flow as a subflow.\n\n``input_slots``/``output_slots`` persist the last-synced subflow interface\n(flow_input/flow_output names in interface order); connections are keyed to\nhandles positionally against them (handle ``input-{i+1}`` <-> input_slots[i];\n``input-0`` is the reserved parameter-data handle).",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"flow_reference": {
"$ref": "#/$defs/SubflowReference"
},
"input_slots": {
"items": {
"type": "string"
},
"title": "Input Slots",
"type": "array"
},
"output_slots": {
"items": {
"type": "string"
},
"title": "Output Slots",
"type": "array"
},
"parameter_specs": {
"items": {
"$ref": "#/$defs/FlowParameter"
},
"title": "Parameter Specs",
"type": "array"
},
"parameter_bindings": {
"items": {
"$ref": "#/$defs/RunFlowParameterBinding"
},
"title": "Parameter Bindings",
"type": "array"
},
"iteration_mode": {
"default": "first_value",
"enum": [
"first_value",
"iterate"
],
"title": "Iteration Mode",
"type": "string"
},
"append_run_metadata": {
"default": true,
"title": "Append Run Metadata",
"type": "boolean"
}
},
"required": [
"flow_id",
"node_id",
"flow_reference"
],
"title": "NodeRunFlow",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
flow_reference(SubflowReference) -
input_slots(list[str]) -
output_slots(list[str]) -
parameter_specs(list[FlowParameter]) -
parameter_bindings(list[RunFlowParameterBinding]) -
iteration_mode(Literal['first_value', 'iterate']) -
append_run_metadata(bool)
Validators:
-
validate_node_reference→node_reference -
_validate_slots_and_bindings
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 | |
input_names
property
Handle labels, index i <-> handle input-i (index 0 = parameter handle).
An empty label at index 0 tells the frontend to hide the parameter handle (the subflow has no parameters); data handles keep input-1..N.
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeSample
pydantic-model
Bases: NodeSingleInput
Settings for a node that samples a subset of the data.
sample_method selects between a cheap top-N slice and a uniform random
sample. It defaults to "first" so flows saved before random sampling
existed (which only carry sample_size) keep their exact behaviour.
fraction is a percentage and only read by "random_fraction";
seed only by the two random methods, where None means a fresh
permutation on every run.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that samples a subset of the data.\n\n``sample_method`` selects between a cheap top-N slice and a uniform random\nsample. It defaults to ``\"first\"`` so flows saved before random sampling\nexisted (which only carry ``sample_size``) keep their exact behaviour.\n``fraction`` is a percentage and only read by ``\"random_fraction\"``;\n``seed`` only by the two random methods, where ``None`` means a fresh\npermutation on every run.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"sample_method": {
"default": "first",
"enum": [
"first",
"random",
"random_fraction"
],
"title": "Sample Method",
"type": "string"
},
"sample_size": {
"default": 1000,
"title": "Sample Size",
"type": "integer"
},
"fraction": {
"default": 10.0,
"title": "Fraction",
"type": "number"
},
"seed": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Seed"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeSample",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
sample_method(SampleMethod) -
sample_size(int) -
fraction(float) -
seed(int | None)
Validators:
-
validate_node_reference→node_reference -
_validate_sample
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 | |
get_default_description()
Describes the sampling method and size.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
640 641 642 643 644 645 646 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeSelect
pydantic-model
Bases: NodeSingleInput
Settings for a node that selects, renames, and reorders columns.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Settings for a node that selects, renames, and reorders columns.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"keep_missing": {
"default": true,
"title": "Keep Missing",
"type": "boolean"
},
"select_input": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Select Input",
"type": "array"
},
"sorted_by": {
"anyOf": [
{
"enum": [
"none",
"asc",
"desc"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "none",
"title": "Sorted By"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeSelect",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
keep_missing(bool) -
select_input(list[SelectInput]) -
sorted_by(Literal['none', 'asc', 'desc'] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 | |
get_default_description()
Describes column selections, renames, and drops.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 | |
to_yaml_dict()
Converts the select node settings to a dictionary for YAML serialization.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeSingleInput
pydantic-model
Bases: NodeBase
A base model for any node that takes a single data input.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "A base model for any node that takes a single data input.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeSingleInput",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
475 476 477 478 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeSort
pydantic-model
Bases: NodeSingleInput
Settings for a node that sorts the data by one or more columns.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SortByInput": {
"description": "Defines a single sort condition on a column, including the direction.",
"properties": {
"column": {
"title": "Column",
"type": "string"
},
"how": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "asc",
"title": "How"
}
},
"required": [
"column"
],
"title": "SortByInput",
"type": "object"
}
},
"description": "Settings for a node that sorts the data by one or more columns.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"sort_input": {
"items": {
"$ref": "#/$defs/SortByInput"
},
"title": "Sort Input",
"type": "array"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeSort",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
sort_input(list[SortByInput])
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
585 586 587 588 589 590 591 592 593 594 595 596 597 598 | |
get_default_description()
Describes the sort columns and directions.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
590 591 592 593 594 595 596 597 598 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeSqlQuery
pydantic-model
Bases: NodeMultiInput
Settings for a node that executes a SQL query against connected data sources.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SqlQueryInput": {
"description": "A container for a SQL query to execute against connected data sources.\n\nNote: ``sql_code`` is *not* validated at schema-construction time. Construction\nis a passive shape-check; the unsafe-SQL gate lives at the executor seam in\n``execute_sql_query`` (and is also enforced by the underlying\n``validate_sql_query`` utility callers can use directly). Validating here too\nwould block legitimate non-AI callers from drafting/testing SQL before\nexecution.",
"properties": {
"sql_code": {
"title": "Sql Code",
"type": "string"
}
},
"required": [
"sql_code"
],
"title": "SqlQueryInput",
"type": "object"
}
},
"description": "Settings for a node that executes a SQL query against connected data sources.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"sql_query_input": {
"$ref": "#/$defs/SqlQueryInput"
}
},
"required": [
"flow_id",
"node_id",
"sql_query_input"
],
"title": "NodeSqlQuery",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
sql_query_input(SqlQueryInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 | |
get_default_description()
Describes the SQL query snippet.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1950 1951 1952 1953 1954 1955 1956 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeTextToRows
pydantic-model
Bases: NodeSingleInput
Settings for a node that splits a text column into multiple rows.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"TextToRowsInput": {
"description": "Defines settings for splitting a text column into multiple rows based on a delimiter.",
"properties": {
"column_to_split": {
"title": "Column To Split",
"type": "string"
},
"output_column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Column Name"
},
"split_by_fixed_value": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Split By Fixed Value"
},
"split_fixed_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": ",",
"title": "Split Fixed Value"
},
"split_by_column": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Split By Column"
}
},
"required": [
"column_to_split"
],
"title": "TextToRowsInput",
"type": "object"
}
},
"description": "Settings for a node that splits a text column into multiple rows.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"text_to_rows_input": {
"$ref": "#/$defs/TextToRowsInput"
}
},
"required": [
"flow_id",
"node_id",
"text_to_rows_input"
],
"title": "NodeTextToRows",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
text_to_rows_input(TextToRowsInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
601 602 603 604 605 606 607 608 609 610 | |
get_default_description()
Describes the text-to-rows split operation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
606 607 608 609 610 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeTrainModel
pydantic-model
Bases: NodeSingleInput
Train an ML model (regression or classification) and optionally publish it to the catalog.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"TrainModelSettings": {
"description": "Settings payload for the Train Model node.\n\n``params`` is a flat dict so the form-driven hyperparameter UI doesn't need\na discriminated union \u2014 the worker validates against the algorithm-specific\nPydantic class via ``shared.ml.trainers.get_trainer(model_type).params_class``.\n\nThe trained model is always written to a flow-scoped path keyed off this\nnode's id so downstream Apply Model nodes in the same flow can read it\nwithout first publishing to the catalog. Set ``publish_to_catalog=True``\nto additionally store the artifact in the catalog (with a stable\ncross-run name + version).",
"properties": {
"target_column": {
"default": "",
"title": "Target Column",
"type": "string"
},
"feature_columns": {
"items": {
"type": "string"
},
"title": "Feature Columns",
"type": "array"
},
"model_type": {
"default": "linear_regression",
"title": "Model Type",
"type": "string"
},
"params": {
"additionalProperties": true,
"title": "Params",
"type": "object"
},
"publish_to_catalog": {
"default": false,
"title": "Publish To Catalog",
"type": "boolean"
},
"model_name": {
"default": "",
"title": "Model Name",
"type": "string"
},
"namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Id"
},
"namespace_full_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Full Name"
},
"catalog_description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Catalog Description"
},
"catalog_tags": {
"items": {
"type": "string"
},
"title": "Catalog Tags",
"type": "array"
}
},
"title": "TrainModelSettings",
"type": "object"
}
},
"description": "Train an ML model (regression or classification) and optionally publish it to the catalog.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"train_input": {
"$ref": "#/$defs/TrainModelSettings"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeTrainModel",
"type": "object"
}
Config:
protected_namespaces:()
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
train_input(TrainModelSettings)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeUnion
pydantic-model
Bases: NodeMultiInput
Settings for a node that concatenates multiple data inputs.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"UnionInput": {
"description": "Defines settings for a union (concatenation) operation.",
"properties": {
"mode": {
"default": "relaxed",
"enum": [
"selective",
"relaxed"
],
"title": "Mode",
"type": "string"
}
},
"title": "UnionInput",
"type": "object"
}
},
"description": "Settings for a node that concatenates multiple data inputs.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"union_input": {
"$ref": "#/$defs/UnionInput"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeUnion",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
union_input(UnionInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1461 1462 1463 1464 1465 1466 1467 1468 | |
get_default_description()
Describes the union mode.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1466 1467 1468 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeUnique
pydantic-model
Bases: NodeSingleInput
Settings for a node that returns the unique rows from the data.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"UniqueInput": {
"description": "Defines settings for a uniqueness operation, specifying columns and which row to keep.",
"properties": {
"columns": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Columns"
},
"strategy": {
"default": "any",
"enum": [
"first",
"last",
"any",
"none"
],
"title": "Strategy",
"type": "string"
}
},
"title": "UniqueInput",
"type": "object"
}
},
"description": "Settings for a node that returns the unique rows from the data.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"unique_input": {
"$ref": "#/$defs/UniqueInput"
}
},
"required": [
"flow_id",
"node_id",
"unique_input"
],
"title": "NodeUnique",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
unique_input(UniqueInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1880 1881 1882 1883 1884 1885 1886 1887 1888 1889 1890 1891 1892 1893 | |
get_default_description()
Describes the uniqueness operation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1885 1886 1887 1888 1889 1890 1891 1892 1893 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeUnpivot
pydantic-model
Bases: NodeSingleInput
Settings for a node that unpivots data from a wide to a long format.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"UnpivotInput": {
"description": "Defines settings for an unpivot (wide-to-long) operation.",
"properties": {
"index_columns": {
"items": {
"type": "string"
},
"title": "Index Columns",
"type": "array"
},
"value_columns": {
"items": {
"type": "string"
},
"title": "Value Columns",
"type": "array"
},
"data_type_selector": {
"anyOf": [
{
"enum": [
"float",
"all",
"date",
"numeric",
"string"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type Selector"
},
"data_type_selector_mode": {
"default": "column",
"enum": [
"data_type",
"column"
],
"title": "Data Type Selector Mode",
"type": "string"
}
},
"title": "UnpivotInput",
"type": "object"
}
},
"description": "Settings for a node that unpivots data from a wide to a long format.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"unpivot_input": {
"$ref": "#/$defs/UnpivotInput",
"default": null
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeUnpivot",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
unpivot_input(UnpivotInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 | |
get_default_description()
Describes the unpivot operation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeWaitFor
pydantic-model
Bases: NodeMultiInput
Pass-through node that enforces ordering on extra dependency inputs.
The first input flows through unchanged; the others have to complete before this node runs but their data is discarded. Useful for enforcing "Apply Model must wait for Train Model" without otherwise coupling their data.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Pass-through node that enforces ordering on extra dependency inputs.\n\nThe first input flows through unchanged; the others have to complete before\nthis node runs but their data is discarded. Useful for enforcing \"Apply\nModel must wait for Train Model\" without otherwise coupling their data.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeWaitFor",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NodeWindowFunctions
pydantic-model
Bases: NodeSingleInput
Settings for a node that adds rolling, cumulative, rank or tile columns.
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
},
"SortByInput": {
"description": "Defines a single sort condition on a column, including the direction.",
"properties": {
"column": {
"title": "Column",
"type": "string"
},
"how": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "asc",
"title": "How"
}
},
"required": [
"column"
],
"title": "SortByInput",
"type": "object"
},
"WindowFunctionInput": {
"description": "A single window-function operation producing one new column.\n\n`column` is the source column for rolling, cumulative and rank functions.\nFor `tile`, `column` is ignored (ordering comes from the outer\n``WindowFunctionsInput.order_by``).",
"properties": {
"column": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Column"
},
"function": {
"enum": [
"rolling_sum",
"rolling_mean",
"rolling_min",
"rolling_max",
"rolling_std",
"cum_sum",
"cum_count",
"cum_min",
"cum_max",
"rank",
"tile"
],
"title": "Function",
"type": "string"
},
"new_column_name": {
"title": "New Column Name",
"type": "string"
},
"window_size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Window Size"
},
"min_periods": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Min Periods"
},
"edge_behavior": {
"anyOf": [
{
"enum": [
"require_full",
"partial",
"fill_zero"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "require_full",
"title": "Edge Behavior"
},
"number_of_groups": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Groups"
},
"rank_method": {
"anyOf": [
{
"enum": [
"ordinal",
"dense",
"min",
"max",
"average"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "ordinal",
"title": "Rank Method"
},
"output_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Type"
}
},
"required": [
"function",
"new_column_name"
],
"title": "WindowFunctionInput",
"type": "object"
},
"WindowFunctionsInput": {
"description": "Defines the settings for a window-functions node.\n\nAttributes\n----------\npartition_by : list[str]\n Optional list of columns to partition by (equivalent to ``.over(...)``).\norder_by : list[SortByInput]\n Ordering within each partition. Required for rolling and tile\n functions; optional (but usually wanted) for cumulative functions.\nwindow_functions : list[WindowFunctionInput]\n Ordered list of per-column window operations to apply. Each produces\n one new column; all are applied in a single ``with_columns`` call.",
"properties": {
"partition_by": {
"items": {
"type": "string"
},
"title": "Partition By",
"type": "array"
},
"order_by": {
"items": {
"$ref": "#/$defs/SortByInput"
},
"title": "Order By",
"type": "array"
},
"window_functions": {
"items": {
"$ref": "#/$defs/WindowFunctionInput"
},
"title": "Window Functions",
"type": "array"
}
},
"title": "WindowFunctionsInput",
"type": "object"
}
},
"description": "Settings for a node that adds rolling, cumulative, rank or tile columns.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": -1,
"title": "Depending On Id"
},
"window_input": {
"$ref": "#/$defs/WindowFunctionsInput"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "NodeWindowFunctions",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_id(int | None) -
window_input(WindowFunctionsInput)
Validators:
-
validate_node_reference→node_reference
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 | |
get_default_description()
Describes the configured window functions.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
NotebookCell
pydantic-model
Bases: BaseModel
A single cell in the notebook editor.
Note: Cell output (stdout, display_outputs, errors) is handled entirely on the frontend and is not persisted. Only id and code are stored.
Show JSON schema:
{
"description": "A single cell in the notebook editor.\n\nNote: Cell output (stdout, display_outputs, errors) is handled entirely\non the frontend and is not persisted. Only id and code are stored.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"code": {
"default": "",
"title": "Code",
"type": "string"
}
},
"required": [
"id"
],
"title": "NotebookCell",
"type": "object"
}
Fields:
-
id(str) -
code(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1959 1960 1961 1962 1963 1964 1965 1966 1967 | |
OutputAvroTable
pydantic-model
Bases: BaseModel
Defines settings for writing an Avro file.
Show JSON schema:
{
"description": "Defines settings for writing an Avro file.",
"properties": {
"file_type": {
"const": "avro",
"default": "avro",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"snappy",
"deflate"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputAvroTable",
"type": "object"
}
Fields:
-
file_type(Literal['avro']) -
compression(Literal['uncompressed', 'snappy', 'deflate'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
318 319 320 321 322 | |
OutputCsvTable
pydantic-model
Bases: BaseModel
Defines settings for writing a CSV file.
Show JSON schema:
{
"description": "Defines settings for writing a CSV file.",
"properties": {
"file_type": {
"const": "csv",
"default": "csv",
"title": "File Type",
"type": "string"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
}
},
"title": "OutputCsvTable",
"type": "object"
}
Fields:
-
file_type(Literal['csv']) -
delimiter(str) -
encoding(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
281 282 283 284 285 286 | |
OutputExcelTable
pydantic-model
Bases: BaseModel
Defines settings for writing an Excel file.
Show JSON schema:
{
"description": "Defines settings for writing an Excel file.",
"properties": {
"file_type": {
"const": "excel",
"default": "excel",
"title": "File Type",
"type": "string"
},
"sheet_name": {
"default": "Sheet1",
"title": "Sheet Name",
"type": "string"
}
},
"title": "OutputExcelTable",
"type": "object"
}
Fields:
-
file_type(Literal['excel']) -
sheet_name(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
297 298 299 300 301 | |
OutputFieldConfig
pydantic-model
Bases: BaseModel
Configuration for output field validation and transformation behavior.
Show JSON schema:
{
"$defs": {
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
}
Fields:
-
enabled(bool) -
validation_mode_behavior(Literal['add_missing', 'add_missing_keep_extra', 'raise_on_missing', 'select_only']) -
fields(list[OutputFieldInfo]) -
validate_data_types(bool)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
97 98 99 100 101 102 103 104 105 106 107 108 | |
OutputFieldInfo
pydantic-model
Bases: BaseModel
Field information with optional default value for output field configuration.
Show JSON schema:
{
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
Fields:
-
name(str) -
data_type(DataTypeStr) -
default_value(str | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
89 90 91 92 93 94 | |
OutputIpcTable
pydantic-model
Bases: BaseModel
Defines settings for writing an Arrow IPC/Feather file.
Show JSON schema:
{
"description": "Defines settings for writing an Arrow IPC/Feather file.",
"properties": {
"file_type": {
"const": "ipc",
"default": "ipc",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"lz4",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputIpcTable",
"type": "object"
}
Fields:
-
file_type(Literal['ipc']) -
compression(Literal['uncompressed', 'lz4', 'zstd'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
304 305 306 307 308 | |
OutputNdjsonTable
pydantic-model
Bases: BaseModel
Defines settings for writing a newline-delimited JSON file.
Show JSON schema:
{
"description": "Defines settings for writing a newline-delimited JSON file.",
"properties": {
"file_type": {
"const": "ndjson",
"default": "ndjson",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"gzip",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputNdjsonTable",
"type": "object"
}
Fields:
-
file_type(Literal['ndjson']) -
compression(Literal['uncompressed', 'gzip', 'zstd'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
311 312 313 314 315 | |
OutputParquetTable
pydantic-model
Bases: BaseModel
Defines settings for writing a Parquet file.
Show JSON schema:
{
"description": "Defines settings for writing a Parquet file.",
"properties": {
"file_type": {
"const": "parquet",
"default": "parquet",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "zstd",
"enum": [
"lz4",
"uncompressed",
"snappy",
"gzip",
"brotli",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputParquetTable",
"type": "object"
}
Fields:
-
file_type(Literal['parquet']) -
compression(Literal['lz4', 'uncompressed', 'snappy', 'gzip', 'brotli', 'zstd'])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
289 290 291 292 293 294 | |
OutputSettings
pydantic-model
Bases: BaseModel
Defines the complete settings for an output node.
Show JSON schema:
{
"$defs": {
"OutputAvroTable": {
"description": "Defines settings for writing an Avro file.",
"properties": {
"file_type": {
"const": "avro",
"default": "avro",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"snappy",
"deflate"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputAvroTable",
"type": "object"
},
"OutputCsvTable": {
"description": "Defines settings for writing a CSV file.",
"properties": {
"file_type": {
"const": "csv",
"default": "csv",
"title": "File Type",
"type": "string"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
}
},
"title": "OutputCsvTable",
"type": "object"
},
"OutputExcelTable": {
"description": "Defines settings for writing an Excel file.",
"properties": {
"file_type": {
"const": "excel",
"default": "excel",
"title": "File Type",
"type": "string"
},
"sheet_name": {
"default": "Sheet1",
"title": "Sheet Name",
"type": "string"
}
},
"title": "OutputExcelTable",
"type": "object"
},
"OutputIpcTable": {
"description": "Defines settings for writing an Arrow IPC/Feather file.",
"properties": {
"file_type": {
"const": "ipc",
"default": "ipc",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"lz4",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputIpcTable",
"type": "object"
},
"OutputNdjsonTable": {
"description": "Defines settings for writing a newline-delimited JSON file.",
"properties": {
"file_type": {
"const": "ndjson",
"default": "ndjson",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "uncompressed",
"enum": [
"uncompressed",
"gzip",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputNdjsonTable",
"type": "object"
},
"OutputParquetTable": {
"description": "Defines settings for writing a Parquet file.",
"properties": {
"file_type": {
"const": "parquet",
"default": "parquet",
"title": "File Type",
"type": "string"
},
"compression": {
"default": "zstd",
"enum": [
"lz4",
"uncompressed",
"snappy",
"gzip",
"brotli",
"zstd"
],
"title": "Compression",
"type": "string"
}
},
"title": "OutputParquetTable",
"type": "object"
}
},
"description": "Defines the complete settings for an output node.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"directory": {
"title": "Directory",
"type": "string"
},
"file_type": {
"title": "File Type",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Fields"
},
"write_mode": {
"default": "overwrite",
"title": "Write Mode",
"type": "string"
},
"table_settings": {
"discriminator": {
"mapping": {
"avro": "#/$defs/OutputAvroTable",
"csv": "#/$defs/OutputCsvTable",
"excel": "#/$defs/OutputExcelTable",
"ipc": "#/$defs/OutputIpcTable",
"ndjson": "#/$defs/OutputNdjsonTable",
"parquet": "#/$defs/OutputParquetTable"
},
"propertyName": "file_type"
},
"oneOf": [
{
"$ref": "#/$defs/OutputCsvTable"
},
{
"$ref": "#/$defs/OutputParquetTable"
},
{
"$ref": "#/$defs/OutputExcelTable"
},
{
"$ref": "#/$defs/OutputIpcTable"
},
{
"$ref": "#/$defs/OutputNdjsonTable"
},
{
"$ref": "#/$defs/OutputAvroTable"
}
],
"title": "Table Settings"
},
"abs_file_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Abs File Path"
}
},
"required": [
"name",
"directory",
"file_type",
"table_settings"
],
"title": "OutputSettings",
"type": "object"
}
Fields:
-
name(str) -
directory(str) -
file_type(str) -
fields(list[str] | None) -
write_mode(str) -
table_settings(OutputTableSettings) -
abs_file_path(str | None)
Validators:
-
validate_table_settings→table_settings -
populate_abs_file_path
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 | |
populate_abs_file_path()
pydantic-validator
Ensures the absolute file path is populated after validation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
420 421 422 423 424 | |
set_absolute_filepath()
Resolves the output directory and name into an absolute path.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
411 412 413 414 415 416 417 418 | |
to_yaml_dict()
Converts the output settings to a dictionary suitable for YAML serialization.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 | |
validate_table_settings(v, info)
pydantic-validator
Ensures table_settings matches the file_type.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 | |
PythonScriptInput
pydantic-model
Bases: BaseModel
Settings for Python code execution on a kernel.
Show JSON schema:
{
"$defs": {
"NotebookCell": {
"description": "A single cell in the notebook editor.\n\nNote: Cell output (stdout, display_outputs, errors) is handled entirely\non the frontend and is not persisted. Only id and code are stored.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"code": {
"default": "",
"title": "Code",
"type": "string"
}
},
"required": [
"id"
],
"title": "NotebookCell",
"type": "object"
}
},
"description": "Settings for Python code execution on a kernel.",
"properties": {
"code": {
"default": "",
"title": "Code",
"type": "string"
},
"kernel_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Kernel Id"
},
"cells": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/NotebookCell"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Cells"
}
},
"title": "PythonScriptInput",
"type": "object"
}
Fields:
-
code(str) -
kernel_id(str | None) -
cells(list[NotebookCell] | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1970 1971 1972 1973 1974 1975 | |
RandomSplitGroup
pydantic-model
Bases: BaseModel
A single output partition in a random split.
Show JSON schema:
{
"description": "A single output partition in a random split.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"percentage": {
"title": "Percentage",
"type": "number"
}
},
"required": [
"name",
"percentage"
],
"title": "RandomSplitGroup",
"type": "object"
}
Fields:
-
name(str) -
percentage(float)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
649 650 651 652 653 | |
RawData
pydantic-model
Bases: BaseModel
Represents data in a raw, columnar format for manual input.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
}
},
"description": "Represents data in a raw, columnar format for manual input.",
"properties": {
"columns": {
"description": "Schema in column order. The i-th MinimalFieldInfo describes the values in data[i].",
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"title": "Columns",
"type": "array"
},
"data": {
"description": "Columnar layout: data[i] is the list of values for columns[i], in column order. len(data) must equal len(columns); each inner list has the same length (one entry per row). For two rows of {name, age}, emit [[\"Alice\", \"Bob\"], [30, 25]] \u2014 NOT [[\"Alice\", 30], [\"Bob\", 25]]. Reading rows back is `data[col_idx][row_idx]`.",
"items": {
"items": {},
"type": "array"
},
"title": "Data",
"type": "array"
}
},
"required": [
"columns",
"data"
],
"title": "RawData",
"type": "object"
}
Fields:
-
columns(list[MinimalFieldInfo]) -
data(list[list])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 | |
columns
pydantic-field
Schema in column order. The i-th MinimalFieldInfo describes the values in data[i].
data
pydantic-field
Columnar layout: data[i] is the list of values for columns[i], in column order. len(data) must equal len(columns); each inner list has the same length (one entry per row). For two rows of {name, age}, emit [["Alice", "Bob"], [30, 25]] — NOT [["Alice", 30], ["Bob", 25]]. Reading rows back is data[col_idx][row_idx].
from_pydict(pydict)
classmethod
Creates a RawData object from a dictionary of lists.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
901 902 903 904 905 906 907 908 909 910 911 | |
from_pylist(pylist)
classmethod
Creates a RawData object from a list of Python dictionaries.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
888 889 890 891 892 893 894 895 896 897 898 899 | |
to_pylist()
Converts the RawData object back into a list of Python dictionaries.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
913 914 915 | |
ReceivedTable
pydantic-model
Bases: BaseModel
Model for defining a table received from an external source.
Show JSON schema:
{
"$defs": {
"InputAvroTable": {
"description": "Defines settings for reading an Avro file.",
"properties": {
"file_type": {
"const": "avro",
"default": "avro",
"title": "File Type",
"type": "string"
}
},
"title": "InputAvroTable",
"type": "object"
},
"InputCsvTable": {
"description": "Defines settings for reading a CSV file.",
"properties": {
"file_type": {
"const": "csv",
"default": "csv",
"title": "File Type",
"type": "string"
},
"reference": {
"default": "",
"title": "Reference",
"type": "string"
},
"starting_from_line": {
"default": 0,
"title": "Starting From Line",
"type": "integer"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
},
"parquet_ref": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Parquet Ref"
},
"row_delimiter": {
"default": "\n",
"title": "Row Delimiter",
"type": "string"
},
"quote_char": {
"default": "\"",
"title": "Quote Char",
"type": "string"
},
"infer_schema_length": {
"default": 10000,
"title": "Infer Schema Length",
"type": "integer"
},
"infer_schema": {
"default": true,
"title": "Infer Schema",
"type": "boolean"
},
"truncate_ragged_lines": {
"default": false,
"title": "Truncate Ragged Lines",
"type": "boolean"
},
"ignore_errors": {
"default": false,
"title": "Ignore Errors",
"type": "boolean"
}
},
"title": "InputCsvTable",
"type": "object"
},
"InputExcelTable": {
"description": "Defines settings for reading an Excel file.",
"properties": {
"file_type": {
"const": "excel",
"default": "excel",
"title": "File Type",
"type": "string"
},
"sheet_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Sheet Name"
},
"start_row": {
"default": 0,
"title": "Start Row",
"type": "integer"
},
"start_column": {
"default": 0,
"title": "Start Column",
"type": "integer"
},
"end_row": {
"default": 0,
"title": "End Row",
"type": "integer"
},
"end_column": {
"default": 0,
"title": "End Column",
"type": "integer"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"type_inference": {
"default": false,
"title": "Type Inference",
"type": "boolean"
}
},
"title": "InputExcelTable",
"type": "object"
},
"InputIpcTable": {
"description": "Defines settings for reading an Arrow IPC/Feather file.",
"properties": {
"file_type": {
"const": "ipc",
"default": "ipc",
"title": "File Type",
"type": "string"
}
},
"title": "InputIpcTable",
"type": "object"
},
"InputJsonTable": {
"description": "Defines settings for reading a JSON file.",
"properties": {
"file_type": {
"const": "json",
"default": "json",
"title": "File Type",
"type": "string"
},
"reference": {
"default": "",
"title": "Reference",
"type": "string"
},
"starting_from_line": {
"default": 0,
"title": "Starting From Line",
"type": "integer"
},
"delimiter": {
"default": ",",
"title": "Delimiter",
"type": "string"
},
"has_headers": {
"default": true,
"title": "Has Headers",
"type": "boolean"
},
"encoding": {
"default": "utf-8",
"title": "Encoding",
"type": "string"
},
"parquet_ref": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Parquet Ref"
},
"row_delimiter": {
"default": "\n",
"title": "Row Delimiter",
"type": "string"
},
"quote_char": {
"default": "\"",
"title": "Quote Char",
"type": "string"
},
"infer_schema_length": {
"default": 10000,
"title": "Infer Schema Length",
"type": "integer"
},
"infer_schema": {
"default": true,
"title": "Infer Schema",
"type": "boolean"
},
"truncate_ragged_lines": {
"default": false,
"title": "Truncate Ragged Lines",
"type": "boolean"
},
"ignore_errors": {
"default": false,
"title": "Ignore Errors",
"type": "boolean"
}
},
"title": "InputJsonTable",
"type": "object"
},
"InputNdjsonTable": {
"description": "Defines settings for reading a newline-delimited JSON file.",
"properties": {
"file_type": {
"const": "ndjson",
"default": "ndjson",
"title": "File Type",
"type": "string"
}
},
"title": "InputNdjsonTable",
"type": "object"
},
"InputParquetTable": {
"description": "Defines settings for reading a Parquet file.",
"properties": {
"file_type": {
"const": "parquet",
"default": "parquet",
"title": "File Type",
"type": "string"
}
},
"title": "InputParquetTable",
"type": "object"
},
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
}
},
"description": "Model for defining a table received from an external source.",
"properties": {
"id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Id"
},
"name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Name"
},
"path": {
"title": "Path",
"type": "string"
},
"directory": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Directory"
},
"analysis_file_available": {
"default": false,
"title": "Analysis File Available",
"type": "boolean"
},
"status": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Status"
},
"fields": {
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"title": "Fields",
"type": "array"
},
"abs_file_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Abs File Path"
},
"file_type": {
"enum": [
"csv",
"json",
"parquet",
"excel",
"ipc",
"ndjson",
"avro"
],
"title": "File Type",
"type": "string"
},
"table_settings": {
"discriminator": {
"mapping": {
"avro": "#/$defs/InputAvroTable",
"csv": "#/$defs/InputCsvTable",
"excel": "#/$defs/InputExcelTable",
"ipc": "#/$defs/InputIpcTable",
"json": "#/$defs/InputJsonTable",
"ndjson": "#/$defs/InputNdjsonTable",
"parquet": "#/$defs/InputParquetTable"
},
"propertyName": "file_type"
},
"oneOf": [
{
"$ref": "#/$defs/InputCsvTable"
},
{
"$ref": "#/$defs/InputJsonTable"
},
{
"$ref": "#/$defs/InputParquetTable"
},
{
"$ref": "#/$defs/InputExcelTable"
},
{
"$ref": "#/$defs/InputIpcTable"
},
{
"$ref": "#/$defs/InputNdjsonTable"
},
{
"$ref": "#/$defs/InputAvroTable"
}
],
"title": "Table Settings"
}
},
"required": [
"path",
"file_type",
"table_settings"
],
"title": "ReceivedTable",
"type": "object"
}
Fields:
-
id(int | None) -
name(str | None) -
path(str) -
directory(str | None) -
analysis_file_available(bool) -
status(str | None) -
fields(list[MinimalFieldInfo]) -
abs_file_path(str | None) -
file_type(Literal['csv', 'json', 'parquet', 'excel', 'ipc', 'ndjson', 'avro']) -
table_settings(InputTableSettings)
Validators:
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 | |
file_path
property
Constructs the full file path from the directory and name.
create_from_path(path, file_type='csv')
classmethod
Creates an instance from a file path string.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 | |
populate_abs_file_path()
pydantic-validator
Ensures the absolute file path is populated after validation.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
273 274 275 276 277 278 | |
set_absolute_filepath()
Resolves the path to an absolute file path.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
249 250 251 252 253 254 255 256 257 258 259 | |
set_default_table_settings(data)
pydantic-validator
Create default table_settings based on file_type if not provided.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
261 262 263 264 265 266 267 268 269 270 271 | |
RemoveItem
pydantic-model
Bases: BaseModel
Represents a single item to be removed from a directory or list.
Show JSON schema:
{
"description": "Represents a single item to be removed from a directory or list.",
"properties": {
"path": {
"title": "Path",
"type": "string"
},
"id": {
"default": -1,
"title": "Id",
"type": "integer"
}
},
"required": [
"path"
],
"title": "RemoveItem",
"type": "object"
}
Fields:
-
path(str) -
id(int)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
68 69 70 71 72 | |
RemoveItemsInput
pydantic-model
Bases: BaseModel
Defines a list of items to be removed.
Show JSON schema:
{
"$defs": {
"RemoveItem": {
"description": "Represents a single item to be removed from a directory or list.",
"properties": {
"path": {
"title": "Path",
"type": "string"
},
"id": {
"default": -1,
"title": "Id",
"type": "integer"
}
},
"required": [
"path"
],
"title": "RemoveItem",
"type": "object"
}
},
"description": "Defines a list of items to be removed.",
"properties": {
"paths": {
"items": {
"$ref": "#/$defs/RemoveItem"
},
"title": "Paths",
"type": "array"
},
"source_path": {
"title": "Source Path",
"type": "string"
}
},
"required": [
"paths",
"source_path"
],
"title": "RemoveItemsInput",
"type": "object"
}
Fields:
-
paths(list[RemoveItem]) -
source_path(str)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
75 76 77 78 79 | |
RestApiAuthSettings
pydantic-model
Bases: BaseModel
Authentication settings for a REST API reader node.
The credential (API key / bearer token / basic password, per auth_type)
is NOT stored inline. secret_name references a secret in the user's
secret store — created once via the Secrets manager and reusable across
nodes — mirroring how the database reader references a stored password. The
.flowfile persists only the reference name, never the credential itself.
secret is an optional inline plaintext for programmatic use
(flowfile_frame.read_api); it is encrypted with the master key and
cleared, never persisted.
Show JSON schema:
{
"description": "Authentication settings for a REST API reader node.\n\nThe credential (API key / bearer token / basic password, per ``auth_type``)\nis NOT stored inline. ``secret_name`` references a secret in the user's\nsecret store \u2014 created once via the Secrets manager and reusable across\nnodes \u2014 mirroring how the database reader references a stored password. The\n``.flowfile`` persists only the reference name, never the credential itself.\n\n``secret`` is an optional inline plaintext for programmatic use\n(``flowfile_frame.read_api``); it is encrypted with the master key and\ncleared, never persisted.",
"properties": {
"auth_type": {
"default": "none",
"enum": [
"none",
"api_key",
"bearer",
"basic"
],
"title": "Auth Type",
"type": "string"
},
"api_key_name": {
"default": "X-API-Key",
"title": "Api Key Name",
"type": "string"
},
"api_key_location": {
"default": "header",
"enum": [
"header",
"query"
],
"title": "Api Key Location",
"type": "string"
},
"basic_username": {
"default": "",
"title": "Basic Username",
"type": "string"
},
"secret_name": {
"default": "",
"title": "Secret Name",
"type": "string"
},
"secret": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Secret"
}
},
"title": "RestApiAuthSettings",
"type": "object"
}
Fields:
-
auth_type(Literal['none', 'api_key', 'bearer', 'basic']) -
api_key_name(str) -
api_key_location(Literal['header', 'query']) -
basic_username(str) -
secret_name(str) -
secret(str | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 | |
RestApiPaginationSettings
pydantic-model
Bases: BaseModel
Pagination strategy and parameters for a REST API reader node.
Show JSON schema:
{
"description": "Pagination strategy and parameters for a REST API reader node.",
"properties": {
"pagination_type": {
"default": "none",
"enum": [
"none",
"offset",
"page",
"cursor"
],
"title": "Pagination Type",
"type": "string"
},
"offset_param": {
"default": "offset",
"title": "Offset Param",
"type": "string"
},
"limit_param": {
"default": "limit",
"title": "Limit Param",
"type": "string"
},
"page_size": {
"default": 100,
"title": "Page Size",
"type": "integer"
},
"page_param": {
"default": "page",
"title": "Page Param",
"type": "string"
},
"start_page": {
"default": 1,
"title": "Start Page",
"type": "integer"
},
"cursor_param": {
"default": "cursor",
"title": "Cursor Param",
"type": "string"
},
"cursor_location": {
"default": "body",
"enum": [
"body",
"header"
],
"title": "Cursor Location",
"type": "string"
},
"cursor_response_path": {
"default": "",
"title": "Cursor Response Path",
"type": "string"
},
"initial_cursor": {
"default": "",
"title": "Initial Cursor",
"type": "string"
},
"max_pages": {
"default": 1000,
"title": "Max Pages",
"type": "integer"
},
"max_records": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Records"
},
"page_delay_seconds": {
"default": 0.0,
"title": "Page Delay Seconds",
"type": "number"
}
},
"title": "RestApiPaginationSettings",
"type": "object"
}
Fields:
-
pagination_type(Literal['none', 'offset', 'page', 'cursor']) -
offset_param(str) -
limit_param(str) -
page_size(int) -
page_param(str) -
start_page(int) -
cursor_param(str) -
cursor_location(Literal['body', 'header']) -
cursor_response_path(str) -
initial_cursor(str) -
max_pages(int) -
max_records(int | None) -
page_delay_seconds(float)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 | |
RestApiSettings
pydantic-model
Bases: BaseModel
UI settings for a REST API reader node.
Secrets are stored inline but encrypted (see RestApiAuthSettings). JSON
is the only supported response format; record_path is a dot-path that
locates the record array within the response body (empty = top-level).
Show JSON schema:
{
"$defs": {
"RestApiAuthSettings": {
"description": "Authentication settings for a REST API reader node.\n\nThe credential (API key / bearer token / basic password, per ``auth_type``)\nis NOT stored inline. ``secret_name`` references a secret in the user's\nsecret store \u2014 created once via the Secrets manager and reusable across\nnodes \u2014 mirroring how the database reader references a stored password. The\n``.flowfile`` persists only the reference name, never the credential itself.\n\n``secret`` is an optional inline plaintext for programmatic use\n(``flowfile_frame.read_api``); it is encrypted with the master key and\ncleared, never persisted.",
"properties": {
"auth_type": {
"default": "none",
"enum": [
"none",
"api_key",
"bearer",
"basic"
],
"title": "Auth Type",
"type": "string"
},
"api_key_name": {
"default": "X-API-Key",
"title": "Api Key Name",
"type": "string"
},
"api_key_location": {
"default": "header",
"enum": [
"header",
"query"
],
"title": "Api Key Location",
"type": "string"
},
"basic_username": {
"default": "",
"title": "Basic Username",
"type": "string"
},
"secret_name": {
"default": "",
"title": "Secret Name",
"type": "string"
},
"secret": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Secret"
}
},
"title": "RestApiAuthSettings",
"type": "object"
},
"RestApiPaginationSettings": {
"description": "Pagination strategy and parameters for a REST API reader node.",
"properties": {
"pagination_type": {
"default": "none",
"enum": [
"none",
"offset",
"page",
"cursor"
],
"title": "Pagination Type",
"type": "string"
},
"offset_param": {
"default": "offset",
"title": "Offset Param",
"type": "string"
},
"limit_param": {
"default": "limit",
"title": "Limit Param",
"type": "string"
},
"page_size": {
"default": 100,
"title": "Page Size",
"type": "integer"
},
"page_param": {
"default": "page",
"title": "Page Param",
"type": "string"
},
"start_page": {
"default": 1,
"title": "Start Page",
"type": "integer"
},
"cursor_param": {
"default": "cursor",
"title": "Cursor Param",
"type": "string"
},
"cursor_location": {
"default": "body",
"enum": [
"body",
"header"
],
"title": "Cursor Location",
"type": "string"
},
"cursor_response_path": {
"default": "",
"title": "Cursor Response Path",
"type": "string"
},
"initial_cursor": {
"default": "",
"title": "Initial Cursor",
"type": "string"
},
"max_pages": {
"default": 1000,
"title": "Max Pages",
"type": "integer"
},
"max_records": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Records"
},
"page_delay_seconds": {
"default": 0.0,
"title": "Page Delay Seconds",
"type": "number"
}
},
"title": "RestApiPaginationSettings",
"type": "object"
}
},
"description": "UI settings for a REST API reader node.\n\nSecrets are stored inline but encrypted (see ``RestApiAuthSettings``). JSON\nis the only supported response format; ``record_path`` is a dot-path that\nlocates the record array within the response body (empty = top-level).",
"properties": {
"url": {
"default": "",
"title": "Url",
"type": "string"
},
"method": {
"default": "GET",
"enum": [
"GET",
"POST"
],
"title": "Method",
"type": "string"
},
"headers": {
"additionalProperties": {
"type": "string"
},
"title": "Headers",
"type": "object"
},
"query_params": {
"additionalProperties": {
"type": "string"
},
"title": "Query Params",
"type": "object"
},
"json_body": {
"anyOf": [
{},
{
"type": "null"
}
],
"default": null,
"title": "Json Body"
},
"auth": {
"$ref": "#/$defs/RestApiAuthSettings"
},
"pagination": {
"$ref": "#/$defs/RestApiPaginationSettings"
},
"record_path": {
"default": "",
"title": "Record Path",
"type": "string"
},
"timeout_seconds": {
"default": 30.0,
"title": "Timeout Seconds",
"type": "number"
},
"max_retries": {
"default": 3,
"title": "Max Retries",
"type": "integer"
}
},
"title": "RestApiSettings",
"type": "object"
}
Fields:
-
url(str) -
method(Literal['GET', 'POST']) -
headers(dict[str, str]) -
query_params(dict[str, str]) -
json_body(Any | None) -
auth(RestApiAuthSettings) -
pagination(RestApiPaginationSettings) -
record_path(str) -
timeout_seconds(float) -
max_retries(int)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 1314 | |
RunFlowParameterBinding
pydantic-model
Bases: BaseModel
How one subflow parameter gets its value for a run_flow execution.
Show JSON schema:
{
"description": "How one subflow parameter gets its value for a run_flow execution.",
"properties": {
"parameter_name": {
"title": "Parameter Name",
"type": "string"
},
"source": {
"default": "default",
"enum": [
"default",
"constant",
"column"
],
"title": "Source",
"type": "string"
},
"constant_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Constant Value"
},
"column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Column Name"
}
},
"required": [
"parameter_name"
],
"title": "RunFlowParameterBinding",
"type": "object"
}
Fields:
-
parameter_name(str) -
source(Literal['default', 'constant', 'column']) -
constant_value(str | None) -
column_name(str | None)
Validators:
-
_validate_source_value
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 | |
SampleUsers
pydantic-model
Bases: ExternalSource
Settings for generating a sample dataset of users.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
}
},
"description": "Settings for generating a sample dataset of users.",
"properties": {
"orientation": {
"default": "row",
"title": "Orientation",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
},
"SAMPLE_USERS": {
"title": "Sample Users",
"type": "boolean"
},
"class_name": {
"default": "sample_users",
"title": "Class Name",
"type": "string"
},
"size": {
"default": 100,
"title": "Size",
"type": "integer"
}
},
"required": [
"SAMPLE_USERS"
],
"title": "SampleUsers",
"type": "object"
}
Fields:
-
orientation(str) -
fields(list[MinimalFieldInfo] | None) -
SAMPLE_USERS(bool) -
class_name(str) -
size(int)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1142 1143 1144 1145 1146 1147 | |
Scd2Settings
pydantic-model
Bases: BaseModel
Slowly-changing-dimension type 2 configuration for a catalog write.
The business key is CatalogWriteSettings.merge_keys — this block only carries the
change-detection scope and the names of the four generated columns. It is persisted verbatim
onto the catalog table record (CatalogTable.scd2_config) so a reader can filter history
without ever reading a writer node's settings.
Show JSON schema:
{
"description": "Slowly-changing-dimension type 2 configuration for a catalog write.\n\nThe business key is ``CatalogWriteSettings.merge_keys`` \u2014 this block only carries the\nchange-detection scope and the names of the four generated columns. It is persisted verbatim\nonto the catalog table record (``CatalogTable.scd2_config``) so a reader can filter history\nwithout ever reading a writer node's settings.",
"properties": {
"compare_columns": {
"items": {
"type": "string"
},
"title": "Compare Columns",
"type": "array"
},
"full_snapshot": {
"default": false,
"title": "Full Snapshot",
"type": "boolean"
},
"partition_on_current": {
"default": true,
"title": "Partition On Current",
"type": "boolean"
},
"surrogate_key_column": {
"default": "sk",
"title": "Surrogate Key Column",
"type": "string"
},
"valid_from_column": {
"default": "valid_from",
"title": "Valid From Column",
"type": "string"
},
"valid_to_column": {
"default": "valid_to",
"title": "Valid To Column",
"type": "string"
},
"is_current_column": {
"default": "is_current",
"title": "Is Current Column",
"type": "string"
}
},
"title": "Scd2Settings",
"type": "object"
}
Fields:
-
compare_columns(list[str]) -
full_snapshot(bool) -
partition_on_current(bool) -
surrogate_key_column(str) -
valid_from_column(str) -
valid_to_column(str) -
is_current_column(str)
Validators:
-
_validate_scd2_columns
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 1657 1658 1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 | |
system_columns
property
The four generated column names, in the primitive's canonical order.
SubflowReference
pydantic-model
Bases: BaseModel
Reference to a catalog-registered flow.
registration_id is the primary reference; flow_uuid is stamped
server-side and used to repair a dangling id; flow_path is display-only.
Show JSON schema:
{
"description": "Reference to a catalog-registered flow.\n\n``registration_id`` is the primary reference; ``flow_uuid`` is stamped\nserver-side and used to repair a dangling id; ``flow_path`` is display-only.",
"properties": {
"registration_id": {
"title": "Registration Id",
"type": "integer"
},
"flow_uuid": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Flow Uuid"
},
"flow_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Flow Path"
}
},
"required": [
"registration_id"
],
"title": "SubflowReference",
"type": "object"
}
Fields:
-
registration_id(int) -
flow_uuid(str | None) -
flow_path(str | None)
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 | |
TrainModelSettings
pydantic-model
Bases: BaseModel
Settings payload for the Train Model node.
params is a flat dict so the form-driven hyperparameter UI doesn't need
a discriminated union — the worker validates against the algorithm-specific
Pydantic class via shared.ml.trainers.get_trainer(model_type).params_class.
The trained model is always written to a flow-scoped path keyed off this
node's id so downstream Apply Model nodes in the same flow can read it
without first publishing to the catalog. Set publish_to_catalog=True
to additionally store the artifact in the catalog (with a stable
cross-run name + version).
Show JSON schema:
{
"description": "Settings payload for the Train Model node.\n\n``params`` is a flat dict so the form-driven hyperparameter UI doesn't need\na discriminated union \u2014 the worker validates against the algorithm-specific\nPydantic class via ``shared.ml.trainers.get_trainer(model_type).params_class``.\n\nThe trained model is always written to a flow-scoped path keyed off this\nnode's id so downstream Apply Model nodes in the same flow can read it\nwithout first publishing to the catalog. Set ``publish_to_catalog=True``\nto additionally store the artifact in the catalog (with a stable\ncross-run name + version).",
"properties": {
"target_column": {
"default": "",
"title": "Target Column",
"type": "string"
},
"feature_columns": {
"items": {
"type": "string"
},
"title": "Feature Columns",
"type": "array"
},
"model_type": {
"default": "linear_regression",
"title": "Model Type",
"type": "string"
},
"params": {
"additionalProperties": true,
"title": "Params",
"type": "object"
},
"publish_to_catalog": {
"default": false,
"title": "Publish To Catalog",
"type": "boolean"
},
"model_name": {
"default": "",
"title": "Model Name",
"type": "string"
},
"namespace_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Id"
},
"namespace_full_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Namespace Full Name"
},
"catalog_description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Catalog Description"
},
"catalog_tags": {
"items": {
"type": "string"
},
"title": "Catalog Tags",
"type": "array"
}
},
"title": "TrainModelSettings",
"type": "object"
}
Config:
protected_namespaces:()
Fields:
-
target_column(str) -
feature_columns(list[str]) -
model_type(str) -
params(dict[str, Any]) -
publish_to_catalog(bool) -
model_name(str) -
namespace_id(int | None) -
namespace_full_name(str | None) -
catalog_description(str | None) -
catalog_tags(list[str])
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 | |
UserDefinedNode
pydantic-model
Bases: NodeMultiInput
Settings for a node that contains the user defined node information
Show JSON schema:
{
"$defs": {
"OutputFieldConfig": {
"description": "Configuration for output field validation and transformation behavior.",
"properties": {
"enabled": {
"default": false,
"title": "Enabled",
"type": "boolean"
},
"validation_mode_behavior": {
"default": "select_only",
"enum": [
"add_missing",
"add_missing_keep_extra",
"raise_on_missing",
"select_only"
],
"title": "Validation Mode Behavior",
"type": "string"
},
"fields": {
"items": {
"$ref": "#/$defs/OutputFieldInfo"
},
"title": "Fields",
"type": "array"
},
"validate_data_types": {
"default": false,
"title": "Validate Data Types",
"type": "boolean"
}
},
"title": "OutputFieldConfig",
"type": "object"
},
"OutputFieldInfo": {
"description": "Field information with optional default value for output field configuration.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"title": "Data Type",
"type": "string"
},
"default_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Default Value"
}
},
"required": [
"name"
],
"title": "OutputFieldInfo",
"type": "object"
}
},
"description": "Settings for a node that contains the user defined node information",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"cache_results": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Cache Results"
},
"pos_x": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos X"
},
"pos_y": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": 0,
"title": "Pos Y"
},
"group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Group Id"
},
"is_setup": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Is Setup"
},
"description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "",
"title": "Description"
},
"node_reference": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Reference"
},
"user_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "User Id"
},
"is_flow_output": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is Flow Output"
},
"is_user_defined": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Is User Defined"
},
"output_field_config": {
"anyOf": [
{
"$ref": "#/$defs/OutputFieldConfig"
},
{
"type": "null"
}
],
"default": null
},
"depending_on_ids": {
"anyOf": [
{
"items": {
"type": "integer"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Depending On Ids"
},
"settings": {
"additionalProperties": {
"additionalProperties": true,
"type": "object"
},
"title": "Settings",
"type": "object"
},
"kernel_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Kernel Id"
},
"output_names": {
"items": {
"type": "string"
},
"title": "Output Names",
"type": "array"
},
"node_source_hash": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Source Hash"
},
"settings_format_version": {
"default": 1,
"title": "Settings Format Version",
"type": "integer"
}
},
"required": [
"flow_id",
"node_id"
],
"title": "UserDefinedNode",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
cache_results(bool | None) -
pos_x(float | None) -
pos_y(float | None) -
group_id(int | None) -
is_setup(bool | None) -
description(str | None) -
node_reference(str | None) -
user_id(int | None) -
is_flow_output(bool | None) -
is_user_defined(bool | None) -
output_field_config(OutputFieldConfig | None) -
depending_on_ids(list[int] | None) -
settings(dict[str, dict[str, Any]]) -
kernel_id(str | None) -
output_names(list[str]) -
node_source_hash(str | None) -
settings_format_version(int)
Validators:
-
validate_node_reference→node_reference -
_coerce_legacy_settings→settings -
validate_output_names→output_names
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 | |
get_default_description()
Generates a human-readable description based on the node's configured content.
Subclasses override this to provide meaningful descriptions. Returns an empty string by default.
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
466 467 468 469 470 471 472 | |
validate_node_reference(v)
pydantic-validator
Validates that node_reference is a safe identifier (lowercase letters, digits, underscores).
Source code in flowfile_core/flowfile_core/schemas/input_schema.py
448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 | |
transform_schema
flowfile_core.schemas.transform_schema
Classes:
| Name | Description |
|---|---|
AggColl |
A data class that represents a single aggregation operation for a group by operation. |
BasicFilter |
Defines a simple, single-condition filter (e.g., 'column' 'equals' 'value'). |
CrossJoinInput |
Data model for cross join operations. |
CrossJoinInputManager |
Manager for cross join operations. |
DynamicRenameInput |
Defines settings for a dynamic rename operation. |
FieldInput |
Represents a single field with its name and data type, typically for defining an output column. |
FilterInput |
Defines the settings for a filter operation, supporting basic or advanced (expression-based) modes. |
FilterOperator |
Supported filter comparison operators. |
FullJoinKeyResponse |
Holds the join key rename responses for both sides of a join. |
FunctionInput |
Defines a formula to be applied, including the output field information. |
FuzzyMatchInput |
Data model for fuzzy matching join operations. |
FuzzyMatchInputManager |
Manager for fuzzy matching join operations. |
GraphSolverInput |
Defines settings for a graph-solving operation (e.g., finding connected components). |
GroupByInput |
A data class that represents the input for a group by operation. |
JoinInput |
Data model for standard SQL-style join operations. |
JoinInputManager |
Manager for standard SQL-style join operations. |
JoinInputs |
Data model for join-specific select inputs (extends SelectInputs). |
JoinInputsManager |
Manager for join-specific operations, extends SelectInputsManager. |
JoinKeyRename |
Represents the renaming of a join key from its original to a temporary name. |
JoinKeyRenameResponse |
Contains a list of join key renames for one side of a join. |
JoinMap |
Defines a single mapping between a left and right column for a join key. |
JoinSelectManagerMixin |
Mixin providing common methods for join-like operations. |
PivotInput |
Defines the settings for a pivot (long-to-wide) operation. |
PolarsCodeInput |
A simple container for a string of user-provided Polars code to be executed. |
RecordIdInput |
Defines settings for adding a record ID (row number) column to the data. |
SelectInput |
Defines how a single column should be selected, renamed, or type-cast. |
SelectInputs |
A container for a list of |
SelectInputsManager |
Manager class that provides all query and mutation operations. |
SortByInput |
Defines a single sort condition on a column, including the direction. |
SqlQueryInput |
A container for a SQL query to execute against connected data sources. |
TextToRowsInput |
Defines settings for splitting a text column into multiple rows based on a delimiter. |
UnionInput |
Defines settings for a union (concatenation) operation. |
UniqueInput |
Defines settings for a uniqueness operation, specifying columns and which row to keep. |
UnpivotInput |
Defines settings for an unpivot (wide-to-long) operation. |
WindowFunctionInput |
A single window-function operation producing one new column. |
WindowFunctionsInput |
Defines the settings for a window-functions node. |
Functions:
| Name | Description |
|---|---|
construct_join_key_name |
Creates a temporary, unique name for a join key column. |
get_func_type_mapping |
Infers the output data type of common aggregation functions. |
get_window_output_type |
Infers the output data type of window functions. |
is_descending |
Whether a sort-direction string means descending. |
string_concat |
A simple wrapper to concatenate string columns in Polars. |
Attributes:
| Name | Type | Description |
|---|---|---|
JoinKeyStrategy |
Key-based join strategies — every option requires |
|
JoinStrategy |
Polars join-strategy enum — broad superset retained for backward |
JoinKeyStrategy = Literal['inner', 'left', 'right', 'full', 'semi', 'anti', 'outer']
module-attribute
Key-based join strategies — every option requires join_mapping
to specify the equality keys. Used by :class:JoinInput.how so the
LLM (and the Pydantic validator) cannot pick "cross" on a
join node — Cartesian joins are the dedicated cross_join
node type's job.
JoinStrategy = Literal['inner', 'left', 'right', 'full', 'semi', 'anti', 'cross', 'outer']
module-attribute
Polars join-strategy enum — broad superset retained for backward
compat with code-generator / flow-data-engine plumbing that still
threads "cross" through DataFrame.join(how="cross") directly.
The join node itself uses :data:JoinKeyStrategy (below) which
excludes "cross" so cross/Cartesian joins route through the
dedicated cross_join node type — making the choice unambiguous
for the AI agent and preventing the join + how="cross" shape
that bypasses the dedicated cross_join node.
AggColl
pydantic-model
Bases: BaseModel
A data class that represents a single aggregation operation for a group by operation.
Attributes
old_name : str The name of the column in the original DataFrame to be aggregated.
str
The aggregation function to use. This can be a string representing a built-in function or a custom function.
Optional[str]
The name of the resulting aggregated column in the output DataFrame. If not provided, it will default to the old_name appended with the aggregation function.
Optional[str]
The type of the output values of the aggregation. If not provided, it is inferred from the aggregation function
using the get_func_type_mapping function.
Example
agg_col = AggColl( old_name='col1', agg='sum', new_name='sum_col1', output_type='float' )
Show JSON schema:
{
"description": "A data class that represents a single aggregation operation for a group by operation.\n\nAttributes\n----------\nold_name : str\n The name of the column in the original DataFrame to be aggregated.\n\nagg : str\n The aggregation function to use. This can be a string representing a built-in function or a custom function.\n\nnew_name : Optional[str]\n The name of the resulting aggregated column in the output DataFrame. If not provided, it will default to the\n old_name appended with the aggregation function.\n\noutput_type : Optional[str]\n The type of the output values of the aggregation. If not provided, it is inferred from the aggregation function\n using the `get_func_type_mapping` function.\n\nExample\n--------\nagg_col = AggColl(\n old_name='col1',\n agg='sum',\n new_name='sum_col1',\n output_type='float'\n)",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"agg": {
"title": "Agg",
"type": "string"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"output_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Type"
}
},
"required": [
"old_name",
"agg"
],
"title": "AggColl",
"type": "object"
}
Fields:
-
old_name(str) -
agg(str) -
new_name(str | None) -
output_type(str | None)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 | |
agg_func
property
Returns the corresponding Polars aggregation function from the agg string.
set_defaults()
pydantic-validator
Set default new_name and output_type based on agg function.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 | |
BasicFilter
pydantic-model
Bases: BaseModel
Defines a simple, single-condition filter (e.g., 'column' 'equals' 'value').
Attributes:
| Name | Type | Description |
|---|---|---|
field |
str
|
The column name to filter on. |
operator |
FilterOperator | str
|
The comparison operator (FilterOperator enum value or symbol). |
value |
str
|
The value to compare against. |
value2 |
str | None
|
Second value for BETWEEN operator (optional). |
Show JSON schema:
{
"$defs": {
"FilterOperator": {
"description": "Supported filter comparison operators.",
"enum": [
"equals",
"not_equals",
"greater_than",
"greater_than_or_equals",
"less_than",
"less_than_or_equals",
"contains",
"not_contains",
"starts_with",
"ends_with",
"is_null",
"is_not_null",
"in",
"not_in",
"between"
],
"title": "FilterOperator",
"type": "string"
}
},
"description": "Defines a simple, single-condition filter (e.g., 'column' 'equals' 'value').\n\nAttributes:\n field: The column name to filter on.\n operator: The comparison operator (FilterOperator enum value or symbol).\n value: The value to compare against.\n value2: Second value for BETWEEN operator (optional).",
"properties": {
"field": {
"default": "",
"title": "Field",
"type": "string"
},
"operator": {
"anyOf": [
{
"$ref": "#/$defs/FilterOperator"
},
{
"type": "string"
}
],
"default": "equals",
"title": "Operator"
},
"value": {
"default": "",
"title": "Value",
"type": "string"
},
"value2": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Value2"
},
"filter_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Type"
},
"filter_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Value"
}
},
"title": "BasicFilter",
"type": "object"
}
Fields:
-
field(str) -
operator(FilterOperator | str) -
value(str) -
value2(str | None) -
filter_type(str | None) -
filter_value(str | None)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 | |
from_yaml_dict(data)
classmethod
Load from YAML format.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
373 374 375 376 377 378 379 380 381 | |
get_operator()
Get the operator as FilterOperator enum.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
356 357 358 359 360 | |
normalize_operator()
pydantic-validator
Normalize the operator to FilterOperator enum.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
345 346 347 348 349 350 351 352 353 354 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
362 363 364 365 366 367 368 369 370 371 | |
CrossJoinInput
pydantic-model
Bases: BaseModel
Data model for cross join operations.
Show JSON schema:
{
"$defs": {
"JoinInputs": {
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Data model for cross join operations.",
"properties": {
"left_select": {
"$ref": "#/$defs/JoinInputs"
},
"right_select": {
"$ref": "#/$defs/JoinInputs"
}
},
"required": [
"left_select",
"right_select"
],
"title": "CrossJoinInput",
"type": "object"
}
Fields:
-
left_select(JoinInputs) -
right_select(JoinInputs)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 | |
__init__(left_select=None, right_select=None, **data)
Custom init for backward compatibility with positional arguments.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
601 602 603 604 605 606 607 608 609 610 611 612 | |
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
621 622 623 624 625 626 | |
parse_inputs(data)
pydantic-validator
Parse flexible input formats before validation.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
614 615 616 617 618 619 | |
CrossJoinInputManager
Bases: JoinSelectManagerMixin
Manager for cross join operations.
Methods:
| Name | Description |
|---|---|
add_new_select_column |
Adds a new column to the selection for either the left or right side. |
auto_generate_new_col_name |
Generates a new, non-conflicting column name by adding a suffix if necessary. |
auto_rename |
Automatically renames columns on the right side to prevent naming conflicts. |
create |
Factory method to create CrossJoinInput from various input formats. |
get_overlapping_columns |
Finds column names that would conflict after the join. |
get_overlapping_records |
Finds column names that would conflict after the join. |
parse_select |
Parses various input formats into a standardized |
to_cross_join_input |
Creates a new CrossJoinInput instance based on the current manager settings. |
Attributes:
| Name | Type | Description |
|---|---|---|
left_select |
JoinInputsManager
|
Backward compatibility: Access left_manager as left_select. |
overlapping_records |
set[str]
|
Backward compatibility: Returns overlapping column names. |
right_select |
JoinInputsManager
|
Backward compatibility: Access right_manager as right_select. |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1394 1395 1396 1397 1398 1399 1400 1401 1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 1412 1413 1414 1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 | |
left_select
property
Backward compatibility: Access left_manager as left_select.
overlapping_records
property
Backward compatibility: Returns overlapping column names.
right_select
property
Backward compatibility: Access right_manager as right_select.
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1385 1386 1387 1388 1389 1390 1391 | |
auto_generate_new_col_name(old_col_name, side)
Generates a new, non-conflicting column name by adding a suffix if necessary.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 | |
auto_rename(rename_mode='prefix')
Automatically renames columns on the right side to prevent naming conflicts.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 | |
create(left_select, right_select)
classmethod
Factory method to create CrossJoinInput from various input formats.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1402 1403 1404 1405 1406 1407 1408 1409 1410 1411 | |
get_overlapping_columns()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1370 1371 1372 | |
get_overlapping_records()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1413 1414 1415 | |
parse_select(select)
staticmethod
Parses various input formats into a standardized JoinInputs object.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 | |
to_cross_join_input()
Creates a new CrossJoinInput instance based on the current manager settings.
This is useful when you've modified the manager (e.g., via auto_rename) and want to get a fresh CrossJoinInput with all the current settings applied.
Returns:
| Type | Description |
|---|---|
CrossJoinInput
|
A new CrossJoinInput instance with current settings |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 | |
DynamicRenameInput
pydantic-model
Bases: BaseModel
Defines settings for a dynamic rename operation.
Applies a single rule (prefix / suffix / formula / first_row) to a set of selected columns, rather than requiring the user to rename columns one-by-one.
In formula mode, the flowfile formula syntax is evaluated with [column_name]
bound to each target column's current name; for example uppercase([column_name])
or "v2_" + [column_name].
In first_row mode, the first row of the incoming table is promoted to column
headers and then dropped from the data. Non-string values are coerced to str;
null or empty values raise an error. Selection filters still apply — only selected
columns are renamed, but the first row is always dropped.
Show JSON schema:
{
"description": "Defines settings for a dynamic rename operation.\n\nApplies a single rule (prefix / suffix / formula / first_row) to a set of selected\ncolumns, rather than requiring the user to rename columns one-by-one.\n\nIn formula mode, the flowfile formula syntax is evaluated with `[column_name]`\nbound to each target column's current name; for example `uppercase([column_name])`\nor `\"v2_\" + [column_name]`.\n\nIn first_row mode, the first row of the incoming table is promoted to column\nheaders and then dropped from the data. Non-string values are coerced to `str`;\nnull or empty values raise an error. Selection filters still apply \u2014 only selected\ncolumns are renamed, but the first row is always dropped.",
"properties": {
"rename_mode": {
"default": "prefix",
"enum": [
"prefix",
"suffix",
"formula",
"first_row"
],
"title": "Rename Mode",
"type": "string"
},
"prefix": {
"default": "",
"title": "Prefix",
"type": "string"
},
"suffix": {
"default": "",
"title": "Suffix",
"type": "string"
},
"formula": {
"default": "",
"expression": true,
"title": "Formula",
"type": "string"
},
"selection_mode": {
"default": "all",
"enum": [
"all",
"list",
"data_type"
],
"title": "Selection Mode",
"type": "string"
},
"selected_columns": {
"items": {
"type": "string"
},
"title": "Selected Columns",
"type": "array"
},
"selected_data_type": {
"anyOf": [
{
"enum": [
"Numeric",
"String",
"Date",
"Other",
"Boolean",
"Binary",
"Complex"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Selected Data Type"
}
},
"title": "DynamicRenameInput",
"type": "object"
}
Fields:
-
rename_mode(RenameMode) -
prefix(str) -
suffix(str) -
formula(str) -
selection_mode(ColumnSelectionMode) -
selected_columns(list[str]) -
selected_data_type(ReadableDataTypeGroup | None)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 | |
FieldInput
pydantic-model
Bases: BaseModel
Represents a single field with its name and data type, typically for defining an output column.
Show JSON schema:
{
"$defs": {
"DataType": {
"description": "Specific data types for fine-grained control.",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Categorical",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array"
],
"title": "DataType",
"type": "string"
}
},
"description": "Represents a single field with its name and data type, typically for defining an output column.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"anyOf": [
{
"$ref": "#/$defs/DataType"
},
{
"const": "Auto",
"type": "string"
},
{
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "Auto",
"title": "Data Type"
}
},
"required": [
"name"
],
"title": "FieldInput",
"type": "object"
}
Fields:
-
name(str) -
data_type(DataType | Literal['Auto'] | DataTypeStr | None)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
276 277 278 279 280 | |
FilterInput
pydantic-model
Bases: BaseModel
Defines the settings for a filter operation, supporting basic or advanced (expression-based) modes.
Attributes:
| Name | Type | Description |
|---|---|---|
mode |
FilterModeLiteral
|
The filter mode - "basic" or "advanced". |
basic_filter |
BasicFilter | None
|
The basic filter configuration (used when mode="basic"). |
advanced_filter |
str
|
The advanced filter expression string (used when mode="advanced"). |
Show JSON schema:
{
"$defs": {
"BasicFilter": {
"description": "Defines a simple, single-condition filter (e.g., 'column' 'equals' 'value').\n\nAttributes:\n field: The column name to filter on.\n operator: The comparison operator (FilterOperator enum value or symbol).\n value: The value to compare against.\n value2: Second value for BETWEEN operator (optional).",
"properties": {
"field": {
"default": "",
"title": "Field",
"type": "string"
},
"operator": {
"anyOf": [
{
"$ref": "#/$defs/FilterOperator"
},
{
"type": "string"
}
],
"default": "equals",
"title": "Operator"
},
"value": {
"default": "",
"title": "Value",
"type": "string"
},
"value2": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Value2"
},
"filter_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Type"
},
"filter_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Value"
}
},
"title": "BasicFilter",
"type": "object"
},
"FilterOperator": {
"description": "Supported filter comparison operators.",
"enum": [
"equals",
"not_equals",
"greater_than",
"greater_than_or_equals",
"less_than",
"less_than_or_equals",
"contains",
"not_contains",
"starts_with",
"ends_with",
"is_null",
"is_not_null",
"in",
"not_in",
"between"
],
"title": "FilterOperator",
"type": "string"
}
},
"description": "Defines the settings for a filter operation, supporting basic or advanced (expression-based) modes.\n\nAttributes:\n mode: The filter mode - \"basic\" or \"advanced\".\n basic_filter: The basic filter configuration (used when mode=\"basic\").\n advanced_filter: The advanced filter expression string (used when mode=\"advanced\").",
"properties": {
"mode": {
"default": "basic",
"enum": [
"basic",
"advanced"
],
"title": "Mode",
"type": "string"
},
"basic_filter": {
"anyOf": [
{
"$ref": "#/$defs/BasicFilter"
},
{
"type": "null"
}
],
"default": null
},
"advanced_filter": {
"default": "",
"expression": true,
"title": "Advanced Filter",
"type": "string"
},
"filter_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Filter Type"
}
},
"title": "FilterInput",
"type": "object"
}
Fields:
-
mode(FilterModeLiteral) -
basic_filter(BasicFilter | None) -
advanced_filter(str) -
filter_type(str | None)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 | |
ensure_basic_filter()
pydantic-validator
Ensure basic_filter exists when mode is basic.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
422 423 424 425 426 427 | |
from_yaml_dict(data)
classmethod
Load from YAML format.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
442 443 444 445 446 447 448 449 450 451 452 453 | |
is_advanced()
Check if filter is in advanced mode.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
429 430 431 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
433 434 435 436 437 438 439 440 | |
FilterOperator
Bases: str, Enum
Supported filter comparison operators.
Methods:
| Name | Description |
|---|---|
from_symbol |
Convert UI symbol to FilterOperator enum. |
to_symbol |
Convert FilterOperator to UI-friendly symbol. |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 | |
from_symbol(symbol)
classmethod
Convert UI symbol to FilterOperator enum.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 | |
to_symbol()
Convert FilterOperator to UI-friendly symbol.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 | |
FullJoinKeyResponse
Bases: NamedTuple
Holds the join key rename responses for both sides of a join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
163 164 165 166 167 | |
FunctionInput
pydantic-model
Bases: BaseModel
Defines a formula to be applied, including the output field information.
Show JSON schema:
{
"$defs": {
"DataType": {
"description": "Specific data types for fine-grained control.",
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Categorical",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array"
],
"title": "DataType",
"type": "string"
},
"FieldInput": {
"description": "Represents a single field with its name and data type, typically for defining an output column.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"anyOf": [
{
"$ref": "#/$defs/DataType"
},
{
"const": "Auto",
"type": "string"
},
{
"enum": [
"Int8",
"Int16",
"Int32",
"Int64",
"Int128",
"UInt8",
"UInt16",
"UInt32",
"UInt64",
"UInt128",
"Float16",
"Float32",
"Float64",
"Decimal",
"String",
"Date",
"Datetime",
"Time",
"Duration",
"Boolean",
"Binary",
"List",
"Struct",
"Array",
"Integer",
"Double",
"Utf8"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "Auto",
"title": "Data Type"
}
},
"required": [
"name"
],
"title": "FieldInput",
"type": "object"
}
},
"description": "Defines a formula to be applied, including the output field information.",
"properties": {
"field": {
"$ref": "#/$defs/FieldInput"
},
"function": {
"expression": true,
"title": "Function",
"type": "string"
}
},
"required": [
"field",
"function"
],
"title": "FunctionInput",
"type": "object"
}
Fields:
-
field(FieldInput) -
function(str)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
283 284 285 286 287 288 289 290 291 292 293 294 | |
FuzzyMatchInput
pydantic-model
Bases: BaseModel
Data model for fuzzy matching join operations.
Show JSON schema:
{
"$defs": {
"FuzzyMapping": {
"description": "Represents the configuration for a fuzzy string match between two columns.\n\nThis class defines all the necessary parameters to perform a fuzzy join,\nincluding the columns to match, the specific algorithm to use, and the\nsimilarity threshold required to consider two strings a match.\n\nIt generates a default name for the output score column if one is not\nprovided.\n\nAttributes:\n left_col (str): The name of the column in the left dataframe to join on.\n right_col (str): The name of the column in the right dataframe to join on.\n threshold_score (float): The similarity score threshold required for a\n match, typically on a scale of 0 to 100. Defaults to 80.0.\n fuzzy_type (FuzzyTypeLiteral): The string-matching algorithm to use.\n Defaults to \"levenshtein\".\n perc_unique (float): A parameter that may be used to assess column\n uniqueness before performing a costly fuzzy match. Defaults to 0.0.\n output_column_name (str | None): The name for the new column that will\n contain the calculated fuzzy match score. If None, a name is\n generated automatically in the format 'fuzzy_score_{left_col}_{right_col}'.\n valid (bool): A flag to indicate whether this mapping is active and should\n be used in a join operation. Defaults to True.\n reversed_threshold_score (float): A property that converts the 0-100\n threshold score into a 0.0-1.0 distance score, where 0.0 is a\n perfect match.",
"properties": {
"left_col": {
"title": "Left Col",
"type": "string"
},
"right_col": {
"title": "Right Col",
"type": "string"
},
"threshold_score": {
"default": 80.0,
"title": "Threshold Score",
"type": "number"
},
"fuzzy_type": {
"default": "levenshtein",
"enum": [
"levenshtein",
"jaro",
"jaro_winkler",
"hamming",
"damerau_levenshtein",
"indel"
],
"title": "Fuzzy Type",
"type": "string"
},
"perc_unique": {
"default": 0.0,
"title": "Perc Unique",
"type": "number"
},
"output_column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Column Name"
},
"valid": {
"default": true,
"title": "Valid",
"type": "boolean"
}
},
"required": [
"left_col",
"right_col"
],
"title": "FuzzyMapping",
"type": "object"
},
"JoinInputs": {
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Data model for fuzzy matching join operations.",
"properties": {
"join_mapping": {
"items": {
"$ref": "#/$defs/FuzzyMapping"
},
"title": "Join Mapping",
"type": "array"
},
"left_select": {
"$ref": "#/$defs/JoinInputs"
},
"right_select": {
"$ref": "#/$defs/JoinInputs"
},
"how": {
"default": "inner",
"enum": [
"inner",
"left",
"right",
"full",
"semi",
"anti",
"cross",
"outer"
],
"title": "How",
"type": "string"
},
"aggregate_output": {
"default": false,
"title": "Aggregate Output",
"type": "boolean"
}
},
"required": [
"join_mapping",
"left_select",
"right_select"
],
"title": "FuzzyMatchInput",
"type": "object"
}
Fields:
-
join_mapping(list[FuzzyMapping]) -
left_select(JoinInputs) -
right_select(JoinInputs) -
how(JoinStrategy) -
aggregate_output(bool)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 | |
__init__(left_select=None, right_select=None, **data)
Custom init for backward compatibility with positional arguments.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
753 754 755 756 757 758 759 760 761 762 763 764 765 | |
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
777 778 779 780 781 782 | |
parse_inputs(data)
pydantic-validator
Parse flexible input formats before validation.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
807 808 809 810 811 812 813 814 815 816 817 818 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
767 768 769 770 771 772 773 774 775 | |
FuzzyMatchInputManager
Bases: JoinInputManager
Manager for fuzzy matching join operations.
Methods:
| Name | Description |
|---|---|
add_new_select_column |
Adds a new column to the selection for either the left or right side. |
auto_generate_new_col_name |
Generates a new, non-conflicting column name by adding a suffix if necessary. |
auto_rename |
Automatically renames columns on the right side to prevent naming conflicts. |
create |
Factory method to create FuzzyMatchInput from various input formats. |
get_fuzzy_maps |
Returns the final fuzzy mappings after applying all column renames. |
get_join_key_renames |
Gets the temporary rename mappings for the join keys on both sides. |
get_left_join_keys |
Returns a set of the left-side join key column names. |
get_left_join_keys_list |
Returns an ordered list of the left-side join key column names. |
get_names_for_table_rename |
Gets join mapping with renamed columns applied. |
get_overlapping_columns |
Finds column names that would conflict after the join. |
get_overlapping_records |
Finds column names that would conflict after the join. |
get_right_join_keys |
Returns a set of the right-side join key column names. |
get_right_join_keys_list |
Returns an ordered list of the right-side join key column names. |
get_used_join_mapping |
Returns the final join mapping after applying all renames and transformations. |
parse_fuzz_mapping |
Parses various input formats into a list of FuzzyMapping objects. |
parse_select |
Parses various input formats into a standardized |
set_join_keys |
Marks the |
to_fuzzy_match_input |
Creates a new FuzzyMatchInput instance based on the current manager settings. |
to_join_input |
Creates a new JoinInput instance based on the current manager settings. |
Attributes:
| Name | Type | Description |
|---|---|---|
aggregate_output |
bool
|
Backward compatibility: Access aggregate_output setting. |
fuzzy_maps |
list[FuzzyMapping]
|
Backward compatibility: Returns fuzzy mappings. |
how |
JoinStrategy
|
Backward compatibility: Access join strategy. |
join_mapping |
list[FuzzyMapping]
|
Backward compatibility: Access fuzzy join mapping. |
left_join_keys |
list[str]
|
Backward compatibility: Returns left join keys list. |
left_select |
JoinInputsManager
|
Backward compatibility: Access left_manager as left_select. |
overlapping_records |
set[str]
|
Backward compatibility: Returns overlapping column names. |
right_join_keys |
list[str]
|
Backward compatibility: Returns right join keys list. |
right_select |
JoinInputsManager
|
Backward compatibility: Access right_manager as right_select. |
used_join_mapping |
list[JoinMap]
|
Backward compatibility: Returns used join mapping. |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1659 1660 1661 1662 1663 1664 1665 1666 1667 1668 1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 1740 1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 1771 1772 1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 | |
aggregate_output
property
Backward compatibility: Access aggregate_output setting.
fuzzy_maps
property
Backward compatibility: Returns fuzzy mappings.
how
property
Backward compatibility: Access join strategy.
join_mapping
property
Backward compatibility: Access fuzzy join mapping.
left_join_keys
property
Backward compatibility: Returns left join keys list.
IMPORTANT: Uses the used_join_mapping PROPERTY (not method).
left_select
property
Backward compatibility: Access left_manager as left_select.
This returns the MANAGER, not the data model. Usage: manager.left_select.join_key_selects
overlapping_records
property
Backward compatibility: Returns overlapping column names.
right_join_keys
property
Backward compatibility: Returns right join keys list.
IMPORTANT: Uses the used_join_mapping PROPERTY (not method).
right_select
property
Backward compatibility: Access right_manager as right_select.
This returns the MANAGER, not the data model. Usage: manager.right_select.join_key_selects
used_join_mapping
property
Backward compatibility: Returns used join mapping.
This property is critical - it's used by left_join_keys and right_join_keys.
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1385 1386 1387 1388 1389 1390 1391 | |
auto_generate_new_col_name(old_col_name, side)
Generates a new, non-conflicting column name by adding a suffix if necessary.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 | |
auto_rename()
Automatically renames columns on the right side to prevent naming conflicts.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 | |
create(join_mapping, left_select, right_select, aggregate_output=False, how='inner')
classmethod
Factory method to create FuzzyMatchInput from various input formats.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 1691 1692 1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 | |
get_fuzzy_maps()
Returns the final fuzzy mappings after applying all column renames.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1741 1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 | |
get_join_key_renames(filter_drop=False)
Gets the temporary rename mappings for the join keys on both sides.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1538 1539 1540 1541 1542 | |
get_left_join_keys()
Returns a set of the left-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1507 1508 1509 | |
get_left_join_keys_list()
Returns an ordered list of the left-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1515 1516 1517 | |
get_names_for_table_rename()
Gets join mapping with renamed columns applied.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 | |
get_overlapping_columns()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1370 1371 1372 | |
get_overlapping_records()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1523 1524 1525 | |
get_right_join_keys()
Returns a set of the right-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1511 1512 1513 | |
get_right_join_keys_list()
Returns an ordered list of the right-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1519 1520 1521 | |
get_used_join_mapping()
Returns the final join mapping after applying all renames and transformations.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 | |
parse_fuzz_mapping(fuzz_mapping)
staticmethod
Parses various input formats into a list of FuzzyMapping objects.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 1728 1729 1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 | |
parse_select(select)
staticmethod
Parses various input formats into a standardized JoinInputs object.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 | |
set_join_keys()
Marks the SelectInput objects corresponding to join keys.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 | |
to_fuzzy_match_input()
Creates a new FuzzyMatchInput instance based on the current manager settings.
This is useful when you've modified the manager (e.g., via auto_rename) and want to get a fresh FuzzyMatchInput with all the current settings applied.
Returns:
| Type | Description |
|---|---|
FuzzyMatchInput
|
A new FuzzyMatchInput instance with current settings |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1778 1779 1780 1781 1782 1783 1784 1785 1786 1787 1788 1789 1790 1791 1792 1793 | |
to_join_input()
Creates a new JoinInput instance based on the current manager settings.
This is useful when you've modified the manager (e.g., via auto_rename) and want to get a fresh JoinInput with all the current settings applied.
Returns:
| Type | Description |
|---|---|
JoinInput
|
A new JoinInput instance with current settings |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 | |
GraphSolverInput
pydantic-model
Bases: BaseModel
Defines settings for a graph-solving operation (e.g., finding connected components).
Show JSON schema:
{
"description": "Defines settings for a graph-solving operation (e.g., finding connected components).",
"properties": {
"col_from": {
"title": "Col From",
"type": "string"
},
"col_to": {
"title": "Col To",
"type": "string"
},
"output_column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "graph_group",
"title": "Output Column Name"
}
},
"required": [
"col_from",
"col_to"
],
"title": "GraphSolverInput",
"type": "object"
}
Fields:
-
col_from(str) -
col_to(str) -
output_column_name(str | None)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1140 1141 1142 1143 1144 1145 | |
GroupByInput
pydantic-model
Bases: BaseModel
A data class that represents the input for a group by operation.
Attributes
agg_cols : List[AggColl]
A list of AggColl objects that specify the aggregation operations to perform on the DataFrame columns
after grouping. Each AggColl object should specify the column to be aggregated and the aggregation
function to use.
Example
group_by_input = GroupByInput( agg_cols=[AggColl(old_name='ix', agg='groupby'), AggColl(old_name='groups', agg='groupby'), AggColl(old_name='col1', agg='sum'), AggColl(old_name='col2', agg='mean')] )
Show JSON schema:
{
"$defs": {
"AggColl": {
"description": "A data class that represents a single aggregation operation for a group by operation.\n\nAttributes\n----------\nold_name : str\n The name of the column in the original DataFrame to be aggregated.\n\nagg : str\n The aggregation function to use. This can be a string representing a built-in function or a custom function.\n\nnew_name : Optional[str]\n The name of the resulting aggregated column in the output DataFrame. If not provided, it will default to the\n old_name appended with the aggregation function.\n\noutput_type : Optional[str]\n The type of the output values of the aggregation. If not provided, it is inferred from the aggregation function\n using the `get_func_type_mapping` function.\n\nExample\n--------\nagg_col = AggColl(\n old_name='col1',\n agg='sum',\n new_name='sum_col1',\n output_type='float'\n)",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"agg": {
"title": "Agg",
"type": "string"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"output_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Type"
}
},
"required": [
"old_name",
"agg"
],
"title": "AggColl",
"type": "object"
}
},
"description": "A data class that represents the input for a group by operation.\n\nAttributes\n----------\nagg_cols : List[AggColl]\n A list of `AggColl` objects that specify the aggregation operations to perform on the DataFrame columns\n after grouping. Each `AggColl` object should specify the column to be aggregated and the aggregation\n function to use.\n\nExample\n--------\ngroup_by_input = GroupByInput(\n agg_cols=[AggColl(old_name='ix', agg='groupby'), AggColl(old_name='groups', agg='groupby'),\n AggColl(old_name='col1', agg='sum'), AggColl(old_name='col2', agg='mean')]\n)",
"properties": {
"agg_cols": {
"items": {
"$ref": "#/$defs/AggColl"
},
"title": "Agg Cols",
"type": "array"
}
},
"required": [
"agg_cols"
],
"title": "GroupByInput",
"type": "object"
}
Fields:
-
agg_cols(list[AggColl])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 | |
__init__(agg_cols)
Backwards compatibility implementation
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
913 914 915 | |
JoinInput
pydantic-model
Bases: BaseModel
Data model for standard SQL-style join operations.
Show JSON schema:
{
"$defs": {
"JoinInputs": {
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
},
"JoinMap": {
"description": "Defines a single mapping between a left and right column for a join key.",
"properties": {
"left_col": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Col"
},
"right_col": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Col"
}
},
"title": "JoinMap",
"type": "object"
},
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Data model for standard SQL-style join operations.",
"properties": {
"join_mapping": {
"items": {
"$ref": "#/$defs/JoinMap"
},
"title": "Join Mapping",
"type": "array"
},
"left_select": {
"$ref": "#/$defs/JoinInputs"
},
"right_select": {
"$ref": "#/$defs/JoinInputs"
},
"how": {
"default": "inner",
"enum": [
"inner",
"left",
"right",
"full",
"semi",
"anti",
"outer"
],
"title": "How",
"type": "string"
}
},
"required": [
"join_mapping",
"left_select",
"right_select"
],
"title": "JoinInput",
"type": "object"
}
Fields:
-
join_mapping(list[JoinMap]) -
left_select(JoinInputs) -
right_select(JoinInputs) -
how(JoinKeyStrategy)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 | |
__init__(join_mapping=None, left_select=None, right_select=None, how='inner', **data)
Custom init for backward compatibility with positional arguments.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 | |
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
736 737 738 739 740 741 | |
parse_inputs(data)
pydantic-validator
Parse flexible input formats before validation.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
727 728 729 730 731 732 733 734 | |
JoinInputManager
Bases: JoinSelectManagerMixin
Manager for standard SQL-style join operations.
Methods:
| Name | Description |
|---|---|
add_new_select_column |
Adds a new column to the selection for either the left or right side. |
auto_generate_new_col_name |
Generates a new, non-conflicting column name by adding a suffix if necessary. |
auto_rename |
Automatically renames columns on the right side to prevent naming conflicts. |
create |
Factory method to create JoinInput from various input formats. |
get_join_key_renames |
Gets the temporary rename mappings for the join keys on both sides. |
get_left_join_keys |
Returns a set of the left-side join key column names. |
get_left_join_keys_list |
Returns an ordered list of the left-side join key column names. |
get_names_for_table_rename |
Gets join mapping with renamed columns applied. |
get_overlapping_columns |
Finds column names that would conflict after the join. |
get_overlapping_records |
Finds column names that would conflict after the join. |
get_right_join_keys |
Returns a set of the right-side join key column names. |
get_right_join_keys_list |
Returns an ordered list of the right-side join key column names. |
get_used_join_mapping |
Returns the final join mapping after applying all renames and transformations. |
parse_select |
Parses various input formats into a standardized |
set_join_keys |
Marks the |
to_join_input |
Creates a new JoinInput instance based on the current manager settings. |
Attributes:
| Name | Type | Description |
|---|---|---|
how |
JoinStrategy
|
Backward compatibility: Access join strategy. |
join_mapping |
list[JoinMap]
|
Backward compatibility: Access join mapping. |
left_join_keys |
list[str]
|
Backward compatibility: Returns left join keys list. |
left_select |
JoinInputsManager
|
Backward compatibility: Access left_manager as left_select. |
overlapping_records |
set[str]
|
Backward compatibility: Returns overlapping column names. |
right_join_keys |
list[str]
|
Backward compatibility: Returns right join keys list. |
right_select |
JoinInputsManager
|
Backward compatibility: Access right_manager as right_select. |
used_join_mapping |
list[JoinMap]
|
Backward compatibility: Returns used join mapping. |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1464 1465 1466 1467 1468 1469 1470 1471 1472 1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 1518 1519 1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 1556 1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 1596 1597 1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 1627 1628 1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 1656 | |
how
property
Backward compatibility: Access join strategy.
join_mapping
property
Backward compatibility: Access join mapping.
left_join_keys
property
Backward compatibility: Returns left join keys list.
IMPORTANT: Uses the used_join_mapping PROPERTY (not method).
left_select
property
Backward compatibility: Access left_manager as left_select.
This returns the MANAGER, not the data model. Usage: manager.left_select.join_key_selects
overlapping_records
property
Backward compatibility: Returns overlapping column names.
right_join_keys
property
Backward compatibility: Returns right join keys list.
IMPORTANT: Uses the used_join_mapping PROPERTY (not method).
right_select
property
Backward compatibility: Access right_manager as right_select.
This returns the MANAGER, not the data model. Usage: manager.right_select.join_key_selects
used_join_mapping
property
Backward compatibility: Returns used join mapping.
This property is critical - it's used by left_join_keys and right_join_keys.
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1385 1386 1387 1388 1389 1390 1391 | |
auto_generate_new_col_name(old_col_name, side)
Generates a new, non-conflicting column name by adding a suffix if necessary.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 | |
auto_rename()
Automatically renames columns on the right side to prevent naming conflicts.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1527 1528 1529 1530 1531 1532 1533 1534 1535 1536 | |
create(join_mapping, left_select, right_select, how='inner')
classmethod
Factory method to create JoinInput from various input formats.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 | |
get_join_key_renames(filter_drop=False)
Gets the temporary rename mappings for the join keys on both sides.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1538 1539 1540 1541 1542 | |
get_left_join_keys()
Returns a set of the left-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1507 1508 1509 | |
get_left_join_keys_list()
Returns an ordered list of the left-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1515 1516 1517 | |
get_names_for_table_rename()
Gets join mapping with renamed columns applied.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1544 1545 1546 1547 1548 1549 1550 1551 1552 1553 1554 1555 | |
get_overlapping_columns()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1370 1371 1372 | |
get_overlapping_records()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1523 1524 1525 | |
get_right_join_keys()
Returns a set of the right-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1511 1512 1513 | |
get_right_join_keys_list()
Returns an ordered list of the right-side join key column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1519 1520 1521 | |
get_used_join_mapping()
Returns the final join mapping after applying all renames and transformations.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1557 1558 1559 1560 1561 1562 1563 1564 1565 1566 1567 1568 1569 1570 1571 1572 1573 | |
parse_select(select)
staticmethod
Parses various input formats into a standardized JoinInputs object.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 | |
set_join_keys()
Marks the SelectInput objects corresponding to join keys.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 | |
to_join_input()
Creates a new JoinInput instance based on the current manager settings.
This is useful when you've modified the manager (e.g., via auto_rename) and want to get a fresh JoinInput with all the current settings applied.
Returns:
| Type | Description |
|---|---|
JoinInput
|
A new JoinInput instance with current settings |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 | |
JoinInputs
pydantic-model
Bases: SelectInputs
Data model for join-specific select inputs (extends SelectInputs).
Show JSON schema:
{
"$defs": {
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "Data model for join-specific select inputs (extends SelectInputs).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "JoinInputs",
"type": "object"
}
Fields:
-
renames(list[SelectInput])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
493 494 495 496 497 498 499 500 501 | |
create_from_list(col_list)
classmethod
Creates a SelectInputs object from a simple list of column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
478 479 480 481 | |
create_from_pl_df(df)
classmethod
Creates a SelectInputs object from a Polars DataFrame's columns.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
483 484 485 486 | |
from_yaml_dict(data)
classmethod
Load from slim YAML format. Supports both 'select' (new) and 'renames' (internal).
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
472 473 474 475 476 | |
remove_select_input(old_key)
Removes a SelectInput from the list based on its original name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
488 489 490 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
468 469 470 | |
JoinInputsManager
Bases: SelectInputsManager
Manager for join-specific operations, extends SelectInputsManager.
Methods:
| Name | Description |
|---|---|
__add__ |
Backward compatibility: Support += operator for appending. |
append |
Appends a new SelectInput to the list of renames. |
find_by_new_name |
Find SelectInput by new column name. |
find_by_old_name |
Find SelectInput by original column name. |
get_drop_columns |
Returns a list of SelectInput objects that are marked to be dropped. |
get_join_key_rename_mapping |
Returns a dictionary mapping original join key names to their temporary names. |
get_join_key_renames |
Gets the temporary rename mapping for all join keys on one side of a join. |
get_join_key_selects |
Returns only the |
get_new_cols |
Returns a set of new (renamed) column names to be kept in the selection. |
get_non_jk_drop_columns |
Returns drop columns that are not join keys. |
get_old_cols |
Returns a set of original column names to be kept in the selection. |
get_rename_table |
Generates a dictionary for use in Polars' |
get_select_cols |
Gets a list of original column names to select from the source DataFrame. |
get_select_input_on_new_name |
Backward compatibility alias: Find SelectInput by new column name. |
get_select_input_on_old_name |
Backward compatibility alias: Find SelectInput by original column name. |
has_drop_cols |
Checks if any column is marked to be dropped from the selection. |
remove_select_input |
Removes a SelectInput from the list based on its original name. |
unselect_field |
Marks a field to be dropped from the final selection by setting |
Attributes:
| Name | Type | Description |
|---|---|---|
drop_columns |
list[SelectInput]
|
Backward compatibility: Returns list of columns to drop. |
join_key_selects |
list[SelectInput]
|
Backward compatibility: Returns join key SelectInputs. |
new_cols |
set[str]
|
Backward compatibility: Returns set of new column names. |
non_jk_drop_columns |
list[SelectInput]
|
Backward compatibility: Returns non-join-key columns to drop. |
old_cols |
set[str]
|
Backward compatibility: Returns set of old column names. |
rename_table |
dict[str, str]
|
Backward compatibility: Returns rename table dictionary. |
renames |
list[SelectInput]
|
Backward compatibility: Direct access to renames list. |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1309 1310 1311 1312 1313 1314 1315 1316 1317 1318 1319 1320 1321 1322 1323 1324 1325 1326 1327 1328 1329 1330 1331 1332 1333 1334 1335 1336 1337 1338 1339 1340 | |
drop_columns
property
Backward compatibility: Returns list of columns to drop.
join_key_selects
property
Backward compatibility: Returns join key SelectInputs.
new_cols
property
Backward compatibility: Returns set of new column names.
non_jk_drop_columns
property
Backward compatibility: Returns non-join-key columns to drop.
old_cols
property
Backward compatibility: Returns set of old column names.
rename_table
property
Backward compatibility: Returns rename table dictionary.
renames
property
Backward compatibility: Direct access to renames list.
__add__(other)
Backward compatibility: Support += operator for appending.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1303 1304 1305 1306 | |
append(other)
Appends a new SelectInput to the list of renames.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1249 1250 1251 | |
find_by_new_name(new_name)
Find SelectInput by new column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1243 1244 1245 | |
find_by_old_name(old_name)
Find SelectInput by original column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1239 1240 1241 | |
get_drop_columns()
Returns a list of SelectInput objects that are marked to be dropped.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1231 1232 1233 | |
get_join_key_rename_mapping(side)
Returns a dictionary mapping original join key names to their temporary names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1332 1333 1334 1335 | |
get_join_key_renames(side, filter_drop=False)
Gets the temporary rename mapping for all join keys on one side of a join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1322 1323 1324 1325 1326 1327 1328 1329 1330 | |
get_join_key_selects()
Returns only the SelectInput objects that are marked as join keys.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1318 1319 1320 | |
get_new_cols()
Returns a set of new (renamed) column names to be kept in the selection.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1211 1212 1213 | |
get_non_jk_drop_columns()
Returns drop columns that are not join keys.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1235 1236 1237 | |
get_old_cols()
Returns a set of original column names to be kept in the selection.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1207 1208 1209 | |
get_rename_table()
Generates a dictionary for use in Polars' .rename() method.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1215 1216 1217 | |
get_select_cols(include_join_key=True)
Gets a list of original column names to select from the source DataFrame.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1219 1220 1221 1222 1223 1224 1225 | |
get_select_input_on_new_name(new_name)
Backward compatibility alias: Find SelectInput by new column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1299 1300 1301 | |
get_select_input_on_old_name(old_name)
Backward compatibility alias: Find SelectInput by original column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1295 1296 1297 | |
has_drop_cols()
Checks if any column is marked to be dropped from the selection.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1227 1228 1229 | |
remove_select_input(old_key)
Removes a SelectInput from the list based on its original name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1253 1254 1255 | |
unselect_field(old_key)
Marks a field to be dropped from the final selection by setting keep to False.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1257 1258 1259 1260 1261 | |
JoinKeyRename
Bases: NamedTuple
Represents the renaming of a join key from its original to a temporary name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
149 150 151 152 153 | |
JoinKeyRenameResponse
Bases: NamedTuple
Contains a list of join key renames for one side of a join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
156 157 158 159 160 | |
JoinMap
pydantic-model
Bases: BaseModel
Defines a single mapping between a left and right column for a join key.
Show JSON schema:
{
"description": "Defines a single mapping between a left and right column for a join key.",
"properties": {
"left_col": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Left Col"
},
"right_col": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Right Col"
}
},
"title": "JoinMap",
"type": "object"
}
Fields:
-
left_col(str | None) -
right_col(str | None)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 | |
set_default_right_col()
pydantic-validator
If right_col is None, default it to left_col.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
517 518 519 520 521 522 | |
JoinSelectManagerMixin
Mixin providing common methods for join-like operations.
Methods:
| Name | Description |
|---|---|
add_new_select_column |
Adds a new column to the selection for either the left or right side. |
auto_generate_new_col_name |
Generates a new, non-conflicting column name by adding a suffix if necessary. |
get_overlapping_columns |
Finds column names that would conflict after the join. |
parse_select |
Parses various input formats into a standardized |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1343 1344 1345 1346 1347 1348 1349 1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 1369 1370 1371 1372 1373 1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 1384 1385 1386 1387 1388 1389 1390 1391 | |
add_new_select_column(select_input, side)
Adds a new column to the selection for either the left or right side.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1385 1386 1387 1388 1389 1390 1391 | |
auto_generate_new_col_name(old_col_name, side)
Generates a new, non-conflicting column name by adding a suffix if necessary.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1374 1375 1376 1377 1378 1379 1380 1381 1382 1383 | |
get_overlapping_columns()
Finds column names that would conflict after the join.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1370 1371 1372 | |
parse_select(select)
staticmethod
Parses various input formats into a standardized JoinInputs object.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1350 1351 1352 1353 1354 1355 1356 1357 1358 1359 1360 1361 1362 1363 1364 1365 1366 1367 1368 | |
PivotInput
pydantic-model
Bases: BaseModel
Defines the settings for a pivot (long-to-wide) operation.
Show JSON schema:
{
"description": "Defines the settings for a pivot (long-to-wide) operation.",
"properties": {
"index_columns": {
"items": {
"type": "string"
},
"title": "Index Columns",
"type": "array"
},
"pivot_column": {
"title": "Pivot Column",
"type": "string"
},
"value_col": {
"title": "Value Col",
"type": "string"
},
"aggregations": {
"items": {
"type": "string"
},
"title": "Aggregations",
"type": "array"
}
},
"required": [
"index_columns",
"pivot_column",
"value_col",
"aggregations"
],
"title": "PivotInput",
"type": "object"
}
Fields:
-
index_columns(list[str]) -
pivot_column(str) -
value_col(str) -
aggregations(list[str])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 | |
grouped_columns
property
Returns the list of columns to be used for the initial grouping stage of the pivot.
get_group_by_input()
Constructs the GroupByInput needed for the pre-aggregation step of the pivot.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
931 932 933 934 935 936 937 | |
get_index_columns()
Returns the index columns as Polars column expressions.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
939 940 941 | |
get_pivot_column()
Returns the pivot column as a Polars column expression.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
943 944 945 | |
get_values_expr()
Creates the struct expression used to gather the values for pivoting.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
947 948 949 | |
PolarsCodeInput
pydantic-model
Bases: BaseModel
A simple container for a string of user-provided Polars code to be executed.
Show JSON schema:
{
"description": "A simple container for a string of user-provided Polars code to be executed.",
"properties": {
"polars_code": {
"title": "Polars Code",
"type": "string"
}
},
"required": [
"polars_code"
],
"title": "PolarsCodeInput",
"type": "object"
}
Fields:
-
polars_code(str)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1179 1180 1181 1182 | |
RecordIdInput
pydantic-model
Bases: BaseModel
Defines settings for adding a record ID (row number) column to the data.
Show JSON schema:
{
"description": "Defines settings for adding a record ID (row number) column to the data.",
"properties": {
"output_column_name": {
"default": "record_id",
"title": "Output Column Name",
"type": "string"
},
"offset": {
"default": 1,
"title": "Offset",
"type": "integer"
},
"group_by": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": false,
"title": "Group By"
},
"group_by_columns": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"title": "Group By Columns"
}
},
"title": "RecordIdInput",
"type": "object"
}
Fields:
-
output_column_name(str) -
offset(int) -
group_by(bool | None) -
group_by_columns(list[str] | None)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
973 974 975 976 977 978 979 | |
SelectInput
pydantic-model
Bases: BaseModel
Defines how a single column should be selected, renamed, or type-cast.
This is a core building block for any operation that involves column manipulation. It holds all the configuration for a single field in a selection operation.
Show JSON schema:
{
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
Config:
frozen:False
Fields:
-
old_name(str) -
original_position(int | None) -
new_name(str | None) -
data_type(str | None) -
data_type_change(bool) -
join_key(bool) -
is_altered(bool) -
position(int | None) -
is_available(bool) -
keep(bool)
Validators:
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 | |
polars_type
property
Translates a user-friendly type name to a Polars data type string.
__eq__(other)
Required when implementing hash.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
257 258 259 260 261 | |
__hash__()
Allow SelectInput to be used in sets and as dict keys.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
253 254 255 | |
from_yaml_dict(data)
classmethod
Load from slim YAML format.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 | |
infer_data_type_change(data)
pydantic-validator
Infer data_type_change when loading from YAML.
When data_type is present but data_type_change is not explicitly set, infer that the user explicitly set the data_type (e.g., when loading from YAML). This ensures is_altered will be set correctly in the after validator.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
228 229 230 231 232 233 234 235 236 237 238 239 240 | |
set_default_new_name()
pydantic-validator
If new_name is None, default it to old_name. Also set is_altered if needed.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
242 243 244 245 246 247 248 249 250 251 | |
to_yaml_dict()
Serialize for YAML output - only user-relevant fields.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
197 198 199 200 201 202 203 204 205 206 207 208 | |
SelectInputs
pydantic-model
Bases: BaseModel
A container for a list of SelectInput objects (pure data, no logic).
Show JSON schema:
{
"$defs": {
"SelectInput": {
"description": "Defines how a single column should be selected, renamed, or type-cast.\n\nThis is a core building block for any operation that involves column manipulation.\nIt holds all the configuration for a single field in a selection operation.",
"properties": {
"old_name": {
"title": "Old Name",
"type": "string"
},
"original_position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Original Position"
},
"new_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "New Name"
},
"data_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type"
},
"data_type_change": {
"default": false,
"title": "Data Type Change",
"type": "boolean"
},
"join_key": {
"default": false,
"title": "Join Key",
"type": "boolean"
},
"is_altered": {
"default": false,
"title": "Is Altered",
"type": "boolean"
},
"position": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Position"
},
"is_available": {
"default": true,
"title": "Is Available",
"type": "boolean"
},
"keep": {
"default": true,
"title": "Keep",
"type": "boolean"
}
},
"required": [
"old_name"
],
"title": "SelectInput",
"type": "object"
}
},
"description": "A container for a list of `SelectInput` objects (pure data, no logic).",
"properties": {
"renames": {
"items": {
"$ref": "#/$defs/SelectInput"
},
"title": "Renames",
"type": "array"
}
},
"title": "SelectInputs",
"type": "object"
}
Fields:
-
renames(list[SelectInput])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 | |
create_from_list(col_list)
classmethod
Creates a SelectInputs object from a simple list of column names.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
478 479 480 481 | |
create_from_pl_df(df)
classmethod
Creates a SelectInputs object from a Polars DataFrame's columns.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
483 484 485 486 | |
from_yaml_dict(data)
classmethod
Load from slim YAML format. Supports both 'select' (new) and 'renames' (internal).
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
472 473 474 475 476 | |
remove_select_input(old_key)
Removes a SelectInput from the list based on its original name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
488 489 490 | |
to_yaml_dict()
Serialize for YAML output.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
468 469 470 | |
SelectInputsManager
Manager class that provides all query and mutation operations.
Methods:
| Name | Description |
|---|---|
__add__ |
Backward compatibility: Support += operator for appending. |
append |
Appends a new SelectInput to the list of renames. |
find_by_new_name |
Find SelectInput by new column name. |
find_by_old_name |
Find SelectInput by original column name. |
get_drop_columns |
Returns a list of SelectInput objects that are marked to be dropped. |
get_new_cols |
Returns a set of new (renamed) column names to be kept in the selection. |
get_non_jk_drop_columns |
Returns drop columns that are not join keys. |
get_old_cols |
Returns a set of original column names to be kept in the selection. |
get_rename_table |
Generates a dictionary for use in Polars' |
get_select_cols |
Gets a list of original column names to select from the source DataFrame. |
get_select_input_on_new_name |
Backward compatibility alias: Find SelectInput by new column name. |
get_select_input_on_old_name |
Backward compatibility alias: Find SelectInput by original column name. |
has_drop_cols |
Checks if any column is marked to be dropped from the selection. |
remove_select_input |
Removes a SelectInput from the list based on its original name. |
unselect_field |
Marks a field to be dropped from the final selection by setting |
Attributes:
| Name | Type | Description |
|---|---|---|
drop_columns |
list[SelectInput]
|
Backward compatibility: Returns list of columns to drop. |
new_cols |
set[str]
|
Backward compatibility: Returns set of new column names. |
non_jk_drop_columns |
list[SelectInput]
|
Backward compatibility: Returns non-join-key columns to drop. |
old_cols |
set[str]
|
Backward compatibility: Returns set of old column names. |
rename_table |
dict[str, str]
|
Backward compatibility: Returns rename table dictionary. |
renames |
list[SelectInput]
|
Backward compatibility: Direct access to renames list. |
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 1245 1246 1247 1248 1249 1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 1266 1267 1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 1284 1285 1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 1300 1301 1302 1303 1304 1305 1306 | |
drop_columns
property
Backward compatibility: Returns list of columns to drop.
new_cols
property
Backward compatibility: Returns set of new column names.
non_jk_drop_columns
property
Backward compatibility: Returns non-join-key columns to drop.
old_cols
property
Backward compatibility: Returns set of old column names.
rename_table
property
Backward compatibility: Returns rename table dictionary.
renames
property
Backward compatibility: Direct access to renames list.
__add__(other)
Backward compatibility: Support += operator for appending.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1303 1304 1305 1306 | |
append(other)
Appends a new SelectInput to the list of renames.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1249 1250 1251 | |
find_by_new_name(new_name)
Find SelectInput by new column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1243 1244 1245 | |
find_by_old_name(old_name)
Find SelectInput by original column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1239 1240 1241 | |
get_drop_columns()
Returns a list of SelectInput objects that are marked to be dropped.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1231 1232 1233 | |
get_new_cols()
Returns a set of new (renamed) column names to be kept in the selection.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1211 1212 1213 | |
get_non_jk_drop_columns()
Returns drop columns that are not join keys.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1235 1236 1237 | |
get_old_cols()
Returns a set of original column names to be kept in the selection.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1207 1208 1209 | |
get_rename_table()
Generates a dictionary for use in Polars' .rename() method.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1215 1216 1217 | |
get_select_cols(include_join_key=True)
Gets a list of original column names to select from the source DataFrame.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1219 1220 1221 1222 1223 1224 1225 | |
get_select_input_on_new_name(new_name)
Backward compatibility alias: Find SelectInput by new column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1299 1300 1301 | |
get_select_input_on_old_name(old_name)
Backward compatibility alias: Find SelectInput by original column name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1295 1296 1297 | |
has_drop_cols()
Checks if any column is marked to be dropped from the selection.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1227 1228 1229 | |
remove_select_input(old_key)
Removes a SelectInput from the list based on its original name.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1253 1254 1255 | |
unselect_field(old_key)
Marks a field to be dropped from the final selection by setting keep to False.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1257 1258 1259 1260 1261 | |
SortByInput
pydantic-model
Bases: BaseModel
Defines a single sort condition on a column, including the direction.
Show JSON schema:
{
"description": "Defines a single sort condition on a column, including the direction.",
"properties": {
"column": {
"title": "Column",
"type": "string"
},
"how": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "asc",
"title": "How"
}
},
"required": [
"column"
],
"title": "SortByInput",
"type": "object"
}
Fields:
-
column(str) -
how(str | None)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
961 962 963 964 965 966 967 968 969 970 | |
descending
property
Resolve how to the boolean Polars expects, accepting both conventions.
SqlQueryInput
pydantic-model
Bases: BaseModel
A container for a SQL query to execute against connected data sources.
Note: sql_code is not validated at schema-construction time. Construction
is a passive shape-check; the unsafe-SQL gate lives at the executor seam in
execute_sql_query (and is also enforced by the underlying
validate_sql_query utility callers can use directly). Validating here too
would block legitimate non-AI callers from drafting/testing SQL before
execution.
Show JSON schema:
{
"description": "A container for a SQL query to execute against connected data sources.\n\nNote: ``sql_code`` is *not* validated at schema-construction time. Construction\nis a passive shape-check; the unsafe-SQL gate lives at the executor seam in\n``execute_sql_query`` (and is also enforced by the underlying\n``validate_sql_query`` utility callers can use directly). Validating here too\nwould block legitimate non-AI callers from drafting/testing SQL before\nexecution.",
"properties": {
"sql_code": {
"title": "Sql Code",
"type": "string"
}
},
"required": [
"sql_code"
],
"title": "SqlQueryInput",
"type": "object"
}
Fields:
-
sql_code(str)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 | |
TextToRowsInput
pydantic-model
Bases: BaseModel
Defines settings for splitting a text column into multiple rows based on a delimiter.
Show JSON schema:
{
"description": "Defines settings for splitting a text column into multiple rows based on a delimiter.",
"properties": {
"column_to_split": {
"title": "Column To Split",
"type": "string"
},
"output_column_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Column Name"
},
"split_by_fixed_value": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Split By Fixed Value"
},
"split_fixed_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": ",",
"title": "Split Fixed Value"
},
"split_by_column": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Split By Column"
}
},
"required": [
"column_to_split"
],
"title": "TextToRowsInput",
"type": "object"
}
Fields:
-
column_to_split(str) -
output_column_name(str | None) -
split_by_fixed_value(bool | None) -
split_fixed_value(str | None) -
split_by_column(str | None)
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1093 1094 1095 1096 1097 1098 1099 1100 | |
UnionInput
pydantic-model
Bases: BaseModel
Defines settings for a union (concatenation) operation.
Show JSON schema:
{
"description": "Defines settings for a union (concatenation) operation.",
"properties": {
"mode": {
"default": "relaxed",
"enum": [
"selective",
"relaxed"
],
"title": "Mode",
"type": "string"
}
},
"title": "UnionInput",
"type": "object"
}
Fields:
-
mode(Literal['selective', 'relaxed'])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1127 1128 1129 1130 | |
UniqueInput
pydantic-model
Bases: BaseModel
Defines settings for a uniqueness operation, specifying columns and which row to keep.
Show JSON schema:
{
"description": "Defines settings for a uniqueness operation, specifying columns and which row to keep.",
"properties": {
"columns": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Columns"
},
"strategy": {
"default": "any",
"enum": [
"first",
"last",
"any",
"none"
],
"title": "Strategy",
"type": "string"
}
},
"title": "UniqueInput",
"type": "object"
}
Fields:
-
columns(list[str] | None) -
strategy(Literal['first', 'last', 'any', 'none'])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1133 1134 1135 1136 1137 | |
UnpivotInput
pydantic-model
Bases: BaseModel
Defines settings for an unpivot (wide-to-long) operation.
Show JSON schema:
{
"description": "Defines settings for an unpivot (wide-to-long) operation.",
"properties": {
"index_columns": {
"items": {
"type": "string"
},
"title": "Index Columns",
"type": "array"
},
"value_columns": {
"items": {
"type": "string"
},
"title": "Value Columns",
"type": "array"
},
"data_type_selector": {
"anyOf": [
{
"enum": [
"float",
"all",
"date",
"numeric",
"string"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Data Type Selector"
},
"data_type_selector_mode": {
"default": "column",
"enum": [
"data_type",
"column"
],
"title": "Data Type Selector Mode",
"type": "string"
}
},
"title": "UnpivotInput",
"type": "object"
}
Config:
arbitrary_types_allowed:True
Fields:
-
index_columns(list[str]) -
value_columns(list[str]) -
data_type_selector(Literal['float', 'all', 'date', 'numeric', 'string'] | None) -
data_type_selector_mode(Literal['data_type', 'column'])
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 | |
data_type_selector_expr
property
Returns a Polars selector function based on the data_type_selector string.
WindowFunctionInput
pydantic-model
Bases: BaseModel
A single window-function operation producing one new column.
column is the source column for rolling, cumulative and rank functions.
For tile, column is ignored (ordering comes from the outer
WindowFunctionsInput.order_by).
Show JSON schema:
{
"description": "A single window-function operation producing one new column.\n\n`column` is the source column for rolling, cumulative and rank functions.\nFor `tile`, `column` is ignored (ordering comes from the outer\n``WindowFunctionsInput.order_by``).",
"properties": {
"column": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Column"
},
"function": {
"enum": [
"rolling_sum",
"rolling_mean",
"rolling_min",
"rolling_max",
"rolling_std",
"cum_sum",
"cum_count",
"cum_min",
"cum_max",
"rank",
"tile"
],
"title": "Function",
"type": "string"
},
"new_column_name": {
"title": "New Column Name",
"type": "string"
},
"window_size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Window Size"
},
"min_periods": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Min Periods"
},
"edge_behavior": {
"anyOf": [
{
"enum": [
"require_full",
"partial",
"fill_zero"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "require_full",
"title": "Edge Behavior"
},
"number_of_groups": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Groups"
},
"rank_method": {
"anyOf": [
{
"enum": [
"ordinal",
"dense",
"min",
"max",
"average"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "ordinal",
"title": "Rank Method"
},
"output_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Type"
}
},
"required": [
"function",
"new_column_name"
],
"title": "WindowFunctionInput",
"type": "object"
}
Fields:
-
column(str | None) -
function(WindowFunctionName) -
new_column_name(str) -
window_size(int | None) -
min_periods(int | None) -
edge_behavior(RollingEdgeBehavior | None) -
number_of_groups(int | None) -
rank_method(RankMethod | None) -
output_type(str | None)
Validators:
-
_validate
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 | |
WindowFunctionsInput
pydantic-model
Bases: BaseModel
Defines the settings for a window-functions node.
Attributes
partition_by : list[str]
Optional list of columns to partition by (equivalent to .over(...)).
order_by : list[SortByInput]
Ordering within each partition. Required for rolling and tile
functions; optional (but usually wanted) for cumulative functions.
window_functions : list[WindowFunctionInput]
Ordered list of per-column window operations to apply. Each produces
one new column; all are applied in a single with_columns call.
Show JSON schema:
{
"$defs": {
"SortByInput": {
"description": "Defines a single sort condition on a column, including the direction.",
"properties": {
"column": {
"title": "Column",
"type": "string"
},
"how": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "asc",
"title": "How"
}
},
"required": [
"column"
],
"title": "SortByInput",
"type": "object"
},
"WindowFunctionInput": {
"description": "A single window-function operation producing one new column.\n\n`column` is the source column for rolling, cumulative and rank functions.\nFor `tile`, `column` is ignored (ordering comes from the outer\n``WindowFunctionsInput.order_by``).",
"properties": {
"column": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Column"
},
"function": {
"enum": [
"rolling_sum",
"rolling_mean",
"rolling_min",
"rolling_max",
"rolling_std",
"cum_sum",
"cum_count",
"cum_min",
"cum_max",
"rank",
"tile"
],
"title": "Function",
"type": "string"
},
"new_column_name": {
"title": "New Column Name",
"type": "string"
},
"window_size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Window Size"
},
"min_periods": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Min Periods"
},
"edge_behavior": {
"anyOf": [
{
"enum": [
"require_full",
"partial",
"fill_zero"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "require_full",
"title": "Edge Behavior"
},
"number_of_groups": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Groups"
},
"rank_method": {
"anyOf": [
{
"enum": [
"ordinal",
"dense",
"min",
"max",
"average"
],
"type": "string"
},
{
"type": "null"
}
],
"default": "ordinal",
"title": "Rank Method"
},
"output_type": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Output Type"
}
},
"required": [
"function",
"new_column_name"
],
"title": "WindowFunctionInput",
"type": "object"
}
},
"description": "Defines the settings for a window-functions node.\n\nAttributes\n----------\npartition_by : list[str]\n Optional list of columns to partition by (equivalent to ``.over(...)``).\norder_by : list[SortByInput]\n Ordering within each partition. Required for rolling and tile\n functions; optional (but usually wanted) for cumulative functions.\nwindow_functions : list[WindowFunctionInput]\n Ordered list of per-column window operations to apply. Each produces\n one new column; all are applied in a single ``with_columns`` call.",
"properties": {
"partition_by": {
"items": {
"type": "string"
},
"title": "Partition By",
"type": "array"
},
"order_by": {
"items": {
"$ref": "#/$defs/SortByInput"
},
"title": "Order By",
"type": "array"
},
"window_functions": {
"items": {
"$ref": "#/$defs/WindowFunctionInput"
},
"title": "Window Functions",
"type": "array"
}
},
"title": "WindowFunctionsInput",
"type": "object"
}
Fields:
-
partition_by(list[str]) -
order_by(list[SortByInput]) -
window_functions(list[WindowFunctionInput])
Validators:
-
_validate
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 | |
construct_join_key_name(side, column_name)
Creates a temporary, unique name for a join key column.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
144 145 146 | |
get_func_type_mapping(func)
Infers the output data type of common aggregation functions.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
105 106 107 108 109 110 111 112 113 114 | |
get_window_output_type(func, input_type=None)
Infers the output data type of window functions.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 | |
is_descending(how)
Whether a sort-direction string means descending.
Accepts both the programmatic "asc"/"desc" form (flowfile_frame) and the
visual editor's "Ascending"/"Descending" form (case-insensitive).
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
952 953 954 955 956 957 958 | |
string_concat(*column)
A simple wrapper to concatenate string columns in Polars.
Source code in flowfile_core/flowfile_core/schemas/transform_schema.py
117 118 119 | |
cloud_storage_schemas
flowfile_core.schemas.cloud_storage_schemas
Cloud storage connection schemas for S3, ADLS, and other cloud providers.
Classes:
| Name | Description |
|---|---|
AuthSettingsInput |
The information needed for the user to provide the details that are needed to provide how to connect to the |
CloudStorageReadSettings |
Settings for reading from cloud storage |
CloudStorageSettings |
Settings for cloud storage nodes in the visual designer |
CloudStorageWriteSettings |
Settings for writing to cloud storage |
CloudStorageWriteSettingsWorkerInterface |
Settings for writing to cloud storage in worker context |
FullCloudStorageConnection |
Internal model with decrypted secrets |
FullCloudStorageConnectionInterface |
API response model - no secrets exposed |
FullCloudStorageConnectionWorkerInterface |
Internal model with decrypted secrets |
WriteSettingsWorkerInterface |
Settings for writing to cloud storage |
Functions:
| Name | Description |
|---|---|
encrypt_for_worker |
Encrypts a secret value for use in worker contexts using per-user key derivation. |
get_cloud_storage_write_settings_worker_interface |
Convert to a worker interface model with encrypted secrets. |
AuthSettingsInput
pydantic-model
Bases: BaseModel
The information needed for the user to provide the details that are needed to provide how to connect to the Cloud provider
Show JSON schema:
{
"description": "The information needed for the user to provide the details that are needed to provide how to connect to the\n Cloud provider",
"properties": {
"storage_type": {
"enum": [
"s3",
"adls",
"gcs"
],
"title": "Storage Type",
"type": "string"
},
"auth_method": {
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account"
],
"title": "Auth Method",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "None",
"title": "Connection Name"
}
},
"required": [
"storage_type",
"auth_method"
],
"title": "AuthSettingsInput",
"type": "object"
}
Fields:
-
storage_type(CloudStorageType) -
auth_method(AuthMethod) -
connection_name(str | None)
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
56 57 58 59 60 61 62 63 64 | |
CloudStorageReadSettings
pydantic-model
Bases: CloudStorageSettings
Settings for reading from cloud storage
Show JSON schema:
{
"description": "Settings for reading from cloud storage",
"properties": {
"auth_mode": {
"default": "auto",
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account",
"auto"
],
"title": "Auth Mode",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Connection Name"
},
"resource_path": {
"title": "Resource Path",
"type": "string"
},
"scan_mode": {
"default": "single_file",
"enum": [
"single_file",
"directory"
],
"title": "Scan Mode",
"type": "string"
},
"file_format": {
"default": "parquet",
"enum": [
"csv",
"parquet",
"json",
"delta",
"iceberg"
],
"title": "File Format",
"type": "string"
},
"csv_has_header": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": true,
"title": "Csv Has Header"
},
"csv_delimiter": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": ",",
"title": "Csv Delimiter"
},
"csv_encoding": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "utf8",
"title": "Csv Encoding"
},
"delta_version": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Delta Version"
}
},
"required": [
"resource_path"
],
"title": "CloudStorageReadSettings",
"type": "object"
}
Fields:
-
auth_mode(CloudStorageAuthMode) -
connection_name(str | None) -
resource_path(str) -
scan_mode(Literal['single_file', 'directory']) -
file_format(Literal['csv', 'parquet', 'json', 'delta', 'iceberg']) -
csv_has_header(bool | None) -
csv_delimiter(str | None) -
csv_encoding(str | None) -
delta_version(int | None)
Validators:
-
validate_auth_requirements→auth_mode
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
188 189 190 191 192 193 194 195 196 | |
CloudStorageSettings
pydantic-model
Bases: BaseModel
Settings for cloud storage nodes in the visual designer
Show JSON schema:
{
"description": "Settings for cloud storage nodes in the visual designer",
"properties": {
"auth_mode": {
"default": "auto",
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account",
"auto"
],
"title": "Auth Mode",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Connection Name"
},
"resource_path": {
"title": "Resource Path",
"type": "string"
}
},
"required": [
"resource_path"
],
"title": "CloudStorageSettings",
"type": "object"
}
Fields:
-
auth_mode(CloudStorageAuthMode) -
connection_name(str | None) -
resource_path(str)
Validators:
-
validate_auth_requirements→auth_mode
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
173 174 175 176 177 178 179 180 181 182 183 184 185 | |
CloudStorageWriteSettings
pydantic-model
Bases: CloudStorageSettings, WriteSettingsWorkerInterface
Settings for writing to cloud storage
Show JSON schema:
{
"description": "Settings for writing to cloud storage",
"properties": {
"resource_path": {
"title": "Resource Path",
"type": "string"
},
"write_mode": {
"default": "overwrite",
"enum": [
"overwrite",
"append"
],
"title": "Write Mode",
"type": "string"
},
"file_format": {
"default": "parquet",
"enum": [
"csv",
"parquet",
"json",
"delta"
],
"title": "File Format",
"type": "string"
},
"parquet_compression": {
"default": "snappy",
"enum": [
"snappy",
"gzip",
"brotli",
"lz4",
"zstd"
],
"title": "Parquet Compression",
"type": "string"
},
"csv_delimiter": {
"default": ",",
"title": "Csv Delimiter",
"type": "string"
},
"csv_encoding": {
"default": "utf8",
"title": "Csv Encoding",
"type": "string"
},
"partition_by": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Partition By"
},
"auth_mode": {
"default": "auto",
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account",
"auto"
],
"title": "Auth Mode",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Connection Name"
}
},
"required": [
"resource_path"
],
"title": "CloudStorageWriteSettings",
"type": "object"
}
Fields:
-
resource_path(str) -
write_mode(Literal['overwrite', 'append']) -
file_format(Literal['csv', 'parquet', 'json', 'delta']) -
parquet_compression(Literal['snappy', 'gzip', 'brotli', 'lz4', 'zstd']) -
csv_delimiter(str) -
csv_encoding(str) -
partition_by(list[str] | None) -
auth_mode(CloudStorageAuthMode) -
connection_name(str | None)
Validators:
-
validate_auth_requirements→auth_mode
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 | |
get_write_setting_worker_interface()
Convert to a worker interface model without secrets.
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
226 227 228 229 230 231 232 233 234 235 236 237 238 | |
CloudStorageWriteSettingsWorkerInterface
pydantic-model
Bases: BaseModel
Settings for writing to cloud storage in worker context
Show JSON schema:
{
"$defs": {
"FullCloudStorageConnectionWorkerInterface": {
"description": "Internal model with decrypted secrets",
"properties": {
"storage_type": {
"enum": [
"s3",
"adls",
"gcs"
],
"title": "Storage Type",
"type": "string"
},
"auth_method": {
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account"
],
"title": "Auth Method",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "None",
"title": "Connection Name"
},
"aws_region": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Region"
},
"aws_access_key_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Access Key Id"
},
"aws_secret_access_key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Secret Access Key"
},
"aws_role_arn": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Role Arn"
},
"aws_allow_unsafe_html": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Allow Unsafe Html"
},
"aws_session_token": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Session Token"
},
"azure_account_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Name"
},
"azure_account_key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Key"
},
"azure_tenant_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Tenant Id"
},
"azure_client_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Id"
},
"azure_client_secret": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Secret"
},
"azure_sas_token": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Sas Token"
},
"gcs_service_account_key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Service Account Key"
},
"gcs_project_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Project Id"
},
"endpoint_url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Endpoint Url"
},
"verify_ssl": {
"default": true,
"title": "Verify Ssl",
"type": "boolean"
}
},
"required": [
"storage_type",
"auth_method"
],
"title": "FullCloudStorageConnectionWorkerInterface",
"type": "object"
},
"WriteSettingsWorkerInterface": {
"description": "Settings for writing to cloud storage",
"properties": {
"resource_path": {
"title": "Resource Path",
"type": "string"
},
"write_mode": {
"default": "overwrite",
"enum": [
"overwrite",
"append"
],
"title": "Write Mode",
"type": "string"
},
"file_format": {
"default": "parquet",
"enum": [
"csv",
"parquet",
"json",
"delta"
],
"title": "File Format",
"type": "string"
},
"parquet_compression": {
"default": "snappy",
"enum": [
"snappy",
"gzip",
"brotli",
"lz4",
"zstd"
],
"title": "Parquet Compression",
"type": "string"
},
"csv_delimiter": {
"default": ",",
"title": "Csv Delimiter",
"type": "string"
},
"csv_encoding": {
"default": "utf8",
"title": "Csv Encoding",
"type": "string"
},
"partition_by": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Partition By"
}
},
"required": [
"resource_path"
],
"title": "WriteSettingsWorkerInterface",
"type": "object"
}
},
"description": "Settings for writing to cloud storage in worker context",
"properties": {
"operation": {
"title": "Operation",
"type": "string"
},
"write_settings": {
"$ref": "#/$defs/WriteSettingsWorkerInterface"
},
"connection": {
"$ref": "#/$defs/FullCloudStorageConnectionWorkerInterface"
},
"flowfile_flow_id": {
"default": 1,
"title": "Flowfile Flow Id",
"type": "integer"
},
"flowfile_node_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "string"
}
],
"default": -1,
"title": "Flowfile Node Id"
}
},
"required": [
"operation",
"write_settings",
"connection"
],
"title": "CloudStorageWriteSettingsWorkerInterface",
"type": "object"
}
Fields:
-
operation(str) -
write_settings(WriteSettingsWorkerInterface) -
connection(FullCloudStorageConnectionWorkerInterface) -
flowfile_flow_id(int) -
flowfile_node_id(int | str)
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
246 247 248 249 250 251 252 253 | |
FullCloudStorageConnection
pydantic-model
Bases: AuthSettingsInput
Internal model with decrypted secrets
Show JSON schema:
{
"description": "Internal model with decrypted secrets",
"properties": {
"storage_type": {
"enum": [
"s3",
"adls",
"gcs"
],
"title": "Storage Type",
"type": "string"
},
"auth_method": {
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account"
],
"title": "Auth Method",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "None",
"title": "Connection Name"
},
"aws_region": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Region"
},
"aws_access_key_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Access Key Id"
},
"aws_secret_access_key": {
"anyOf": [
{
"format": "password",
"type": "string",
"writeOnly": true
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Secret Access Key"
},
"aws_role_arn": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Role Arn"
},
"aws_allow_unsafe_html": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Allow Unsafe Html"
},
"aws_session_token": {
"anyOf": [
{
"format": "password",
"type": "string",
"writeOnly": true
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Session Token"
},
"azure_account_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Name"
},
"azure_account_key": {
"anyOf": [
{
"format": "password",
"type": "string",
"writeOnly": true
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Key"
},
"azure_tenant_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Tenant Id"
},
"azure_client_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Id"
},
"azure_client_secret": {
"anyOf": [
{
"format": "password",
"type": "string",
"writeOnly": true
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Secret"
},
"azure_sas_token": {
"anyOf": [
{
"format": "password",
"type": "string",
"writeOnly": true
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Sas Token"
},
"gcs_service_account_key": {
"anyOf": [
{
"format": "password",
"type": "string",
"writeOnly": true
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Service Account Key"
},
"gcs_project_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Project Id"
},
"endpoint_url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Endpoint Url"
},
"verify_ssl": {
"default": true,
"title": "Verify Ssl",
"type": "boolean"
}
},
"required": [
"storage_type",
"auth_method"
],
"title": "FullCloudStorageConnection",
"type": "object"
}
Fields:
-
storage_type(CloudStorageType) -
auth_method(AuthMethod) -
connection_name(str | None) -
aws_region(str | None) -
aws_access_key_id(str | None) -
aws_secret_access_key(SecretStr | None) -
aws_role_arn(str | None) -
aws_allow_unsafe_html(bool | None) -
aws_session_token(SecretStr | None) -
azure_account_name(str | None) -
azure_account_key(SecretStr | None) -
azure_tenant_id(str | None) -
azure_client_id(str | None) -
azure_client_secret(SecretStr | None) -
azure_sas_token(SecretStr | None) -
gcs_service_account_key(SecretStr | None) -
gcs_project_id(str | None) -
endpoint_url(str | None) -
verify_ssl(bool)
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | |
get_worker_interface(user_id)
Convert to a worker interface model with encrypted secrets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
user_id
|
int
|
The user ID for per-user key derivation |
required |
Returns:
| Type | Description |
|---|---|
FullCloudStorageConnectionWorkerInterface
|
FullCloudStorageConnectionWorkerInterface with encrypted secrets |
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | |
FullCloudStorageConnectionInterface
pydantic-model
Bases: AuthSettingsInput
API response model - no secrets exposed
Show JSON schema:
{
"$defs": {
"AccessInfo": {
"description": "How the requesting user can access a resource; attached to list/detail responses.",
"properties": {
"is_owner": {
"title": "Is Owner",
"type": "boolean"
},
"access_level": {
"enum": [
"owner",
"manage",
"use"
],
"title": "Access Level",
"type": "string"
},
"shared_by": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Shared By"
}
},
"required": [
"is_owner",
"access_level"
],
"title": "AccessInfo",
"type": "object"
}
},
"description": "API response model - no secrets exposed",
"properties": {
"storage_type": {
"enum": [
"s3",
"adls",
"gcs"
],
"title": "Storage Type",
"type": "string"
},
"auth_method": {
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account"
],
"title": "Auth Method",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "None",
"title": "Connection Name"
},
"aws_allow_unsafe_html": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Allow Unsafe Html"
},
"aws_region": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Region"
},
"aws_access_key_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Access Key Id"
},
"aws_role_arn": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Role Arn"
},
"azure_account_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Name"
},
"azure_tenant_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Tenant Id"
},
"azure_client_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Id"
},
"gcs_project_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Project Id"
},
"endpoint_url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Endpoint Url"
},
"verify_ssl": {
"default": true,
"title": "Verify Ssl",
"type": "boolean"
},
"id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Id"
},
"access": {
"anyOf": [
{
"$ref": "#/$defs/AccessInfo"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"storage_type",
"auth_method"
],
"title": "FullCloudStorageConnectionInterface",
"type": "object"
}
Fields:
-
storage_type(CloudStorageType) -
auth_method(AuthMethod) -
connection_name(str | None) -
aws_allow_unsafe_html(bool | None) -
aws_region(str | None) -
aws_access_key_id(str | None) -
aws_role_arn(str | None) -
azure_account_name(str | None) -
azure_tenant_id(str | None) -
azure_client_id(str | None) -
gcs_project_id(str | None) -
endpoint_url(str | None) -
verify_ssl(bool) -
id(int | None) -
access(AccessInfo | None)
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 | |
FullCloudStorageConnectionWorkerInterface
pydantic-model
Bases: AuthSettingsInput
Internal model with decrypted secrets
Show JSON schema:
{
"description": "Internal model with decrypted secrets",
"properties": {
"storage_type": {
"enum": [
"s3",
"adls",
"gcs"
],
"title": "Storage Type",
"type": "string"
},
"auth_method": {
"enum": [
"access_key",
"iam_role",
"service_principal",
"managed_identity",
"sas_token",
"aws-cli",
"env_vars",
"service_account"
],
"title": "Auth Method",
"type": "string"
},
"connection_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": "None",
"title": "Connection Name"
},
"aws_region": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Region"
},
"aws_access_key_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Access Key Id"
},
"aws_secret_access_key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Secret Access Key"
},
"aws_role_arn": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Role Arn"
},
"aws_allow_unsafe_html": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Allow Unsafe Html"
},
"aws_session_token": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Aws Session Token"
},
"azure_account_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Name"
},
"azure_account_key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Account Key"
},
"azure_tenant_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Tenant Id"
},
"azure_client_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Id"
},
"azure_client_secret": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Client Secret"
},
"azure_sas_token": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Azure Sas Token"
},
"gcs_service_account_key": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Service Account Key"
},
"gcs_project_id": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Gcs Project Id"
},
"endpoint_url": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Endpoint Url"
},
"verify_ssl": {
"default": true,
"title": "Verify Ssl",
"type": "boolean"
}
},
"required": [
"storage_type",
"auth_method"
],
"title": "FullCloudStorageConnectionWorkerInterface",
"type": "object"
}
Fields:
-
storage_type(CloudStorageType) -
auth_method(AuthMethod) -
connection_name(str | None) -
aws_region(str | None) -
aws_access_key_id(str | None) -
aws_secret_access_key(str | None) -
aws_role_arn(str | None) -
aws_allow_unsafe_html(bool | None) -
aws_session_token(str | None) -
azure_account_name(str | None) -
azure_account_key(str | None) -
azure_tenant_id(str | None) -
azure_client_id(str | None) -
azure_client_secret(str | None) -
azure_sas_token(str | None) -
gcs_service_account_key(str | None) -
gcs_project_id(str | None) -
endpoint_url(str | None) -
verify_ssl(bool)
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 | |
WriteSettingsWorkerInterface
pydantic-model
Bases: BaseModel
Settings for writing to cloud storage
Show JSON schema:
{
"description": "Settings for writing to cloud storage",
"properties": {
"resource_path": {
"title": "Resource Path",
"type": "string"
},
"write_mode": {
"default": "overwrite",
"enum": [
"overwrite",
"append"
],
"title": "Write Mode",
"type": "string"
},
"file_format": {
"default": "parquet",
"enum": [
"csv",
"parquet",
"json",
"delta"
],
"title": "File Format",
"type": "string"
},
"parquet_compression": {
"default": "snappy",
"enum": [
"snappy",
"gzip",
"brotli",
"lz4",
"zstd"
],
"title": "Parquet Compression",
"type": "string"
},
"csv_delimiter": {
"default": ",",
"title": "Csv Delimiter",
"type": "string"
},
"csv_encoding": {
"default": "utf8",
"title": "Csv Encoding",
"type": "string"
},
"partition_by": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Partition By"
}
},
"required": [
"resource_path"
],
"title": "WriteSettingsWorkerInterface",
"type": "object"
}
Fields:
-
resource_path(str) -
write_mode(Literal['overwrite', 'append']) -
file_format(Literal['csv', 'parquet', 'json', 'delta']) -
parquet_compression(Literal['snappy', 'gzip', 'brotli', 'lz4', 'zstd']) -
csv_delimiter(str) -
csv_encoding(str) -
partition_by(list[str] | None)
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 | |
encrypt_for_worker(secret_value, user_id)
Encrypts a secret value for use in worker contexts using per-user key derivation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
secret_value
|
SecretStr | None
|
The secret value to encrypt |
required |
user_id
|
int
|
The user ID for key derivation |
required |
Returns:
| Type | Description |
|---|---|
str | None
|
Encrypted secret with embedded user_id, or None if secret_value is None |
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
40 41 42 43 44 45 46 47 48 49 50 51 52 53 | |
get_cloud_storage_write_settings_worker_interface(write_settings, connection, lf, user_id, flowfile_flow_id=1, flowfile_node_id=-1)
Convert to a worker interface model with encrypted secrets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
write_settings
|
CloudStorageWriteSettings
|
Cloud storage write settings |
required |
connection
|
FullCloudStorageConnection
|
Full cloud storage connection with secrets |
required |
lf
|
LazyFrame
|
LazyFrame to serialize |
required |
user_id
|
int
|
User ID for per-user key derivation |
required |
flowfile_flow_id
|
int
|
Flow ID for tracking |
1
|
flowfile_node_id
|
int | str
|
Node ID for tracking |
-1
|
Returns:
| Type | Description |
|---|---|
CloudStorageWriteSettingsWorkerInterface
|
CloudStorageWriteSettingsWorkerInterface ready for worker |
Source code in flowfile_core/flowfile_core/schemas/cloud_storage_schemas.py
256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 | |
output_model
flowfile_core.schemas.output_model
Classes:
| Name | Description |
|---|---|
BaseItem |
A base model for any item in a file system, like a file or directory. |
ExpressionRef |
A reference to a single Polars expression, including its name and docstring. |
ExpressionsOverview |
Represents a categorized list of available Polars expressions. |
FileColumn |
Represents detailed schema and statistics for a single column (field). |
InstantFuncResult |
Represents the result of a function that is expected to execute instantly. |
ItemInfo |
Provides detailed information about a single item in an output directory. |
NodeData |
A comprehensive model holding the complete state and data for a single node. |
NodeDescriptionResponse |
Response model for the node description endpoint. |
NodeInputNameInfo |
Describes a named input available for a kernel node. |
NodeResult |
Represents the execution result of a single node in a FlowGraph run. |
OutputDir |
Represents the contents of a single output directory. |
OutputFile |
Represents a single file in an output directory, extending BaseItem. |
OutputFiles |
Represents a collection of files, typically within a directory. |
OutputTree |
Represents a directory tree, including subdirectories. |
ProjectExportFile |
A single file in a project export (path relative to the project root). |
ProjectExportManifest |
The full file manifest of a flow exported as a Python project. |
ProjectSaveRequest |
Request to write a project export to a directory on the server. |
ProjectSaveResponse |
Result of writing a project export to disk. |
RunInformation |
Contains summary information about a complete FlowGraph execution. |
TableExample |
Represents a preview of a table, including schema and sample data. |
BaseItem
pydantic-model
Bases: BaseModel
A base model for any item in a file system, like a file or directory.
Show JSON schema:
{
"description": "A base model for any item in a file system, like a file or directory.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
}
},
"required": [
"name",
"path"
],
"title": "BaseItem",
"type": "object"
}
Fields:
-
name(str) -
path(str) -
size(int | None) -
creation_date(datetime | None) -
access_date(datetime | None) -
modification_date(datetime | None) -
source_path(str | None) -
number_of_items(int)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
43 44 45 46 47 48 49 50 51 52 53 | |
ExpressionRef
pydantic-model
Bases: BaseModel
A reference to a single Polars expression, including its name and docstring.
Show JSON schema:
{
"description": "A reference to a single Polars expression, including its name and docstring.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"doc": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"title": "Doc"
}
},
"required": [
"name",
"doc"
],
"title": "ExpressionRef",
"type": "object"
}
Fields:
-
name(str) -
doc(str | None)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
162 163 164 165 166 | |
ExpressionsOverview
pydantic-model
Bases: BaseModel
Represents a categorized list of available Polars expressions.
Show JSON schema:
{
"$defs": {
"ExpressionRef": {
"description": "A reference to a single Polars expression, including its name and docstring.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"doc": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"title": "Doc"
}
},
"required": [
"name",
"doc"
],
"title": "ExpressionRef",
"type": "object"
}
},
"description": "Represents a categorized list of available Polars expressions.",
"properties": {
"expression_type": {
"title": "Expression Type",
"type": "string"
},
"expressions": {
"items": {
"$ref": "#/$defs/ExpressionRef"
},
"title": "Expressions",
"type": "array"
}
},
"required": [
"expression_type",
"expressions"
],
"title": "ExpressionsOverview",
"type": "object"
}
Fields:
-
expression_type(str) -
expressions(list[ExpressionRef])
Source code in flowfile_core/flowfile_core/schemas/output_model.py
169 170 171 172 173 | |
FileColumn
pydantic-model
Bases: BaseModel
Represents detailed schema and statistics for a single column (field).
The statistics fields are None until they are actually computed — either
never (plain schema previews) or exactly, on demand, via the column-stats
endpoint writing into the node's FlowfileColumn.
Show JSON schema:
{
"description": "Represents detailed schema and statistics for a single column (field).\n\nThe statistics fields are None until they are actually computed \u2014 either\nnever (plain schema previews) or exactly, on demand, via the column-stats\nendpoint writing into the node's ``FlowfileColumn``.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"title": "Data Type",
"type": "string"
},
"data_type_group": {
"default": "Other",
"enum": [
"Numeric",
"String",
"Date",
"Other",
"Boolean",
"Binary",
"Complex"
],
"title": "Data Type Group",
"type": "string"
},
"is_unique": {
"default": false,
"title": "Is Unique",
"type": "boolean"
},
"max_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Value"
},
"min_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Min Value"
},
"average_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Average Value"
},
"number_of_empty_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Empty Values"
},
"number_of_filled_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Filled Values"
},
"number_of_unique_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Unique Values"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
}
},
"required": [
"name",
"data_type"
],
"title": "FileColumn",
"type": "object"
}
Fields:
-
name(str) -
data_type(str) -
data_type_group(ReadableDataTypeGroup) -
is_unique(bool) -
max_value(str | None) -
min_value(str | None) -
average_value(str | None) -
number_of_empty_values(int | None) -
number_of_filled_values(int | None) -
number_of_unique_values(int | None) -
size(int | None)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 | |
InstantFuncResult
pydantic-model
Bases: BaseModel
Represents the result of a function that is expected to execute instantly.
Show JSON schema:
{
"description": "Represents the result of a function that is expected to execute instantly.",
"properties": {
"success": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Success"
},
"result": {
"title": "Result",
"type": "string"
}
},
"required": [
"result"
],
"title": "InstantFuncResult",
"type": "object"
}
Fields:
-
success(bool | None) -
result(str)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
176 177 178 179 180 | |
ItemInfo
pydantic-model
Bases: OutputFile
Provides detailed information about a single item in an output directory.
Show JSON schema:
{
"description": "Provides detailed information about a single item in an output directory.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"ext": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ext"
},
"mimetype": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Mimetype"
},
"id": {
"default": -1,
"title": "Id",
"type": "integer"
},
"type": {
"title": "Type",
"type": "string"
},
"analysis_file_available": {
"default": false,
"title": "Analysis File Available",
"type": "boolean"
},
"analysis_file_location": {
"default": null,
"title": "Analysis File Location",
"type": "string"
},
"analysis_file_error": {
"default": null,
"title": "Analysis File Error",
"type": "string"
}
},
"required": [
"name",
"path",
"type"
],
"title": "ItemInfo",
"type": "object"
}
Fields:
-
name(str) -
path(str) -
size(int | None) -
creation_date(datetime | None) -
access_date(datetime | None) -
modification_date(datetime | None) -
source_path(str | None) -
number_of_items(int) -
ext(str | None) -
mimetype(str | None) -
id(int) -
type(str) -
analysis_file_available(bool) -
analysis_file_location(str) -
analysis_file_error(str)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
145 146 147 148 149 150 151 152 | |
NodeData
pydantic-model
Bases: BaseModel
A comprehensive model holding the complete state and data for a single node.
This includes its input/output data previews, settings, and run status.
Show JSON schema:
{
"$defs": {
"FileColumn": {
"description": "Represents detailed schema and statistics for a single column (field).\n\nThe statistics fields are None until they are actually computed \u2014 either\nnever (plain schema previews) or exactly, on demand, via the column-stats\nendpoint writing into the node's ``FlowfileColumn``.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"title": "Data Type",
"type": "string"
},
"data_type_group": {
"default": "Other",
"enum": [
"Numeric",
"String",
"Date",
"Other",
"Boolean",
"Binary",
"Complex"
],
"title": "Data Type Group",
"type": "string"
},
"is_unique": {
"default": false,
"title": "Is Unique",
"type": "boolean"
},
"max_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Value"
},
"min_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Min Value"
},
"average_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Average Value"
},
"number_of_empty_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Empty Values"
},
"number_of_filled_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Filled Values"
},
"number_of_unique_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Unique Values"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
}
},
"required": [
"name",
"data_type"
],
"title": "FileColumn",
"type": "object"
},
"TableExample": {
"description": "Represents a preview of a table, including schema and sample data.\n\n``number_of_records`` is None when the total is unknown (e.g. a lazy result\nwhose count was never computed); 0 always means a genuinely empty result.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"number_of_records": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Records"
},
"number_of_columns": {
"title": "Number Of Columns",
"type": "integer"
},
"name": {
"title": "Name",
"type": "string"
},
"table_schema": {
"items": {
"$ref": "#/$defs/FileColumn"
},
"title": "Table Schema",
"type": "array"
},
"columns": {
"items": {
"type": "string"
},
"title": "Columns",
"type": "array"
},
"data": {
"anyOf": [
{
"items": {
"additionalProperties": true,
"type": "object"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Data"
},
"has_example_data": {
"default": false,
"title": "Has Example Data",
"type": "boolean"
},
"has_run_with_current_setup": {
"default": false,
"title": "Has Run With Current Setup",
"type": "boolean"
}
},
"required": [
"node_id",
"number_of_columns",
"name",
"table_schema",
"columns"
],
"title": "TableExample",
"type": "object"
}
},
"description": "A comprehensive model holding the complete state and data for a single node.\n\nThis includes its input/output data previews, settings, and run status.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"node_id": {
"title": "Node Id",
"type": "integer"
},
"flow_type": {
"title": "Flow Type",
"type": "string"
},
"left_input": {
"anyOf": [
{
"$ref": "#/$defs/TableExample"
},
{
"type": "null"
}
],
"default": null
},
"right_input": {
"anyOf": [
{
"$ref": "#/$defs/TableExample"
},
{
"type": "null"
}
],
"default": null
},
"main_input": {
"anyOf": [
{
"$ref": "#/$defs/TableExample"
},
{
"type": "null"
}
],
"default": null
},
"main_output": {
"anyOf": [
{
"$ref": "#/$defs/TableExample"
},
{
"type": "null"
}
],
"default": null
},
"left_output": {
"anyOf": [
{
"$ref": "#/$defs/TableExample"
},
{
"type": "null"
}
],
"default": null
},
"right_output": {
"anyOf": [
{
"$ref": "#/$defs/TableExample"
},
{
"type": "null"
}
],
"default": null
},
"has_run": {
"default": false,
"title": "Has Run",
"type": "boolean"
},
"is_cached": {
"default": false,
"title": "Is Cached",
"type": "boolean"
},
"setting_input": {
"default": null,
"title": "Setting Input"
},
"prediction_warning": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Prediction Warning"
}
},
"required": [
"flow_id",
"node_id",
"flow_type"
],
"title": "NodeData",
"type": "object"
}
Fields:
-
flow_id(int) -
node_id(int) -
flow_type(str) -
left_input(TableExample | None) -
right_input(TableExample | None) -
main_input(TableExample | None) -
main_output(TableExample | None) -
left_output(TableExample | None) -
right_output(TableExample | None) -
has_run(bool) -
is_cached(bool) -
setting_input(Any) -
prediction_warning(str | None)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 | |
NodeDescriptionResponse
pydantic-model
Bases: BaseModel
Response model for the node description endpoint.
Show JSON schema:
{
"description": "Response model for the node description endpoint.",
"properties": {
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"is_auto_generated": {
"default": false,
"title": "Is Auto Generated",
"type": "boolean"
}
},
"title": "NodeDescriptionResponse",
"type": "object"
}
Fields:
-
description(str) -
is_auto_generated(bool)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
183 184 185 186 187 | |
NodeInputNameInfo
pydantic-model
Bases: BaseModel
Describes a named input available for a kernel node.
Show JSON schema:
{
"description": "Describes a named input available for a kernel node.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"source_node_id": {
"title": "Source Node Id",
"type": "integer"
},
"source_node_type": {
"title": "Source Node Type",
"type": "string"
}
},
"required": [
"name",
"source_node_id",
"source_node_type"
],
"title": "NodeInputNameInfo",
"type": "object"
}
Fields:
-
name(str) -
source_node_id(int) -
source_node_type(str)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
95 96 97 98 99 100 | |
NodeResult
pydantic-model
Bases: BaseModel
Represents the execution result of a single node in a FlowGraph run.
Show JSON schema:
{
"description": "Represents the execution result of a single node in a FlowGraph run.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"node_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Name"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"start_timestamp": {
"title": "Start Timestamp",
"type": "number"
},
"end_timestamp": {
"default": 0,
"title": "End Timestamp",
"type": "number"
},
"success": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Success"
},
"error": {
"default": "",
"title": "Error",
"type": "string"
},
"run_time_ms": {
"default": -1,
"description": "Run time in milliseconds",
"title": "Run Time Ms",
"type": "integer"
},
"is_running": {
"default": true,
"title": "Is Running",
"type": "boolean"
}
},
"required": [
"node_id"
],
"title": "NodeResult",
"type": "object"
}
Fields:
-
node_id(int) -
node_name(str | None) -
description(str) -
start_timestamp(float) -
end_timestamp(float) -
success(bool | None) -
error(str) -
run_time_ms(int) -
is_running(bool)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 | |
run_time_ms = -1
pydantic-field
Run time in milliseconds
OutputDir
pydantic-model
Bases: BaseItem
Represents the contents of a single output directory.
Show JSON schema:
{
"$defs": {
"ItemInfo": {
"description": "Provides detailed information about a single item in an output directory.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"ext": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ext"
},
"mimetype": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Mimetype"
},
"id": {
"default": -1,
"title": "Id",
"type": "integer"
},
"type": {
"title": "Type",
"type": "string"
},
"analysis_file_available": {
"default": false,
"title": "Analysis File Available",
"type": "boolean"
},
"analysis_file_location": {
"default": null,
"title": "Analysis File Location",
"type": "string"
},
"analysis_file_error": {
"default": null,
"title": "Analysis File Error",
"type": "string"
}
},
"required": [
"name",
"path",
"type"
],
"title": "ItemInfo",
"type": "object"
}
},
"description": "Represents the contents of a single output directory.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"all_items": {
"items": {
"type": "string"
},
"title": "All Items",
"type": "array"
},
"items": {
"items": {
"$ref": "#/$defs/ItemInfo"
},
"title": "Items",
"type": "array"
}
},
"required": [
"name",
"path",
"all_items",
"items"
],
"title": "OutputDir",
"type": "object"
}
Fields:
-
name(str) -
path(str) -
size(int | None) -
creation_date(datetime | None) -
access_date(datetime | None) -
modification_date(datetime | None) -
source_path(str | None) -
number_of_items(int) -
all_items(list[str]) -
items(list[ItemInfo])
Source code in flowfile_core/flowfile_core/schemas/output_model.py
155 156 157 158 159 | |
OutputFile
pydantic-model
Bases: BaseItem
Represents a single file in an output directory, extending BaseItem.
Show JSON schema:
{
"description": "Represents a single file in an output directory, extending BaseItem.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"ext": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ext"
},
"mimetype": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Mimetype"
}
},
"required": [
"name",
"path"
],
"title": "OutputFile",
"type": "object"
}
Fields:
-
name(str) -
path(str) -
size(int | None) -
creation_date(datetime | None) -
access_date(datetime | None) -
modification_date(datetime | None) -
source_path(str | None) -
number_of_items(int) -
ext(str | None) -
mimetype(str | None)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
126 127 128 129 130 | |
OutputFiles
pydantic-model
Bases: BaseItem
Represents a collection of files, typically within a directory.
Show JSON schema:
{
"$defs": {
"OutputFile": {
"description": "Represents a single file in an output directory, extending BaseItem.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"ext": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ext"
},
"mimetype": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Mimetype"
}
},
"required": [
"name",
"path"
],
"title": "OutputFile",
"type": "object"
}
},
"description": "Represents a collection of files, typically within a directory.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"files": {
"items": {
"$ref": "#/$defs/OutputFile"
},
"title": "Files",
"type": "array"
}
},
"required": [
"name",
"path"
],
"title": "OutputFiles",
"type": "object"
}
Fields:
-
name(str) -
path(str) -
size(int | None) -
creation_date(datetime | None) -
access_date(datetime | None) -
modification_date(datetime | None) -
source_path(str | None) -
number_of_items(int) -
files(list[OutputFile])
Source code in flowfile_core/flowfile_core/schemas/output_model.py
133 134 135 136 | |
OutputTree
pydantic-model
Bases: OutputFiles
Represents a directory tree, including subdirectories.
Show JSON schema:
{
"$defs": {
"OutputFile": {
"description": "Represents a single file in an output directory, extending BaseItem.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"ext": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ext"
},
"mimetype": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Mimetype"
}
},
"required": [
"name",
"path"
],
"title": "OutputFile",
"type": "object"
},
"OutputFiles": {
"description": "Represents a collection of files, typically within a directory.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"files": {
"items": {
"$ref": "#/$defs/OutputFile"
},
"title": "Files",
"type": "array"
}
},
"required": [
"name",
"path"
],
"title": "OutputFiles",
"type": "object"
}
},
"description": "Represents a directory tree, including subdirectories.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"path": {
"title": "Path",
"type": "string"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
},
"creation_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Creation Date"
},
"access_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Access Date"
},
"modification_date": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Modification Date"
},
"source_path": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Source Path"
},
"number_of_items": {
"default": -1,
"title": "Number Of Items",
"type": "integer"
},
"files": {
"items": {
"$ref": "#/$defs/OutputFile"
},
"title": "Files",
"type": "array"
},
"directories": {
"items": {
"$ref": "#/$defs/OutputFiles"
},
"title": "Directories",
"type": "array"
}
},
"required": [
"name",
"path"
],
"title": "OutputTree",
"type": "object"
}
Fields:
-
name(str) -
path(str) -
size(int | None) -
creation_date(datetime | None) -
access_date(datetime | None) -
modification_date(datetime | None) -
source_path(str | None) -
number_of_items(int) -
files(list[OutputFile]) -
directories(list[OutputFiles])
Source code in flowfile_core/flowfile_core/schemas/output_model.py
139 140 141 142 | |
ProjectExportFile
pydantic-model
Bases: BaseModel
A single file in a project export (path relative to the project root).
Show JSON schema:
{
"description": "A single file in a project export (path relative to the project root).",
"properties": {
"path": {
"title": "Path",
"type": "string"
},
"content": {
"title": "Content",
"type": "string"
}
},
"required": [
"path",
"content"
],
"title": "ProjectExportFile",
"type": "object"
}
Fields:
-
path(str) -
content(str)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
190 191 192 193 194 | |
ProjectExportManifest
pydantic-model
Bases: BaseModel
The full file manifest of a flow exported as a Python project.
Show JSON schema:
{
"$defs": {
"ProjectExportFile": {
"description": "A single file in a project export (path relative to the project root).",
"properties": {
"path": {
"title": "Path",
"type": "string"
},
"content": {
"title": "Content",
"type": "string"
}
},
"required": [
"path",
"content"
],
"title": "ProjectExportFile",
"type": "object"
}
},
"description": "The full file manifest of a flow exported as a Python project.",
"properties": {
"project_name": {
"title": "Project Name",
"type": "string"
},
"files": {
"items": {
"$ref": "#/$defs/ProjectExportFile"
},
"title": "Files",
"type": "array"
},
"warnings": {
"items": {
"type": "string"
},
"title": "Warnings",
"type": "array"
}
},
"required": [
"project_name",
"files"
],
"title": "ProjectExportManifest",
"type": "object"
}
Fields:
-
project_name(str) -
files(list[ProjectExportFile]) -
warnings(list[str])
Source code in flowfile_core/flowfile_core/schemas/output_model.py
197 198 199 200 201 202 | |
ProjectSaveRequest
pydantic-model
Bases: BaseModel
Request to write a project export to a directory on the server.
Show JSON schema:
{
"description": "Request to write a project export to a directory on the server.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"target_directory": {
"title": "Target Directory",
"type": "string"
},
"overwrite": {
"default": false,
"title": "Overwrite",
"type": "boolean"
}
},
"required": [
"flow_id",
"target_directory"
],
"title": "ProjectSaveRequest",
"type": "object"
}
Fields:
-
flow_id(int) -
target_directory(str) -
overwrite(bool)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
205 206 207 208 209 210 | |
ProjectSaveResponse
pydantic-model
Bases: BaseModel
Result of writing a project export to disk.
Show JSON schema:
{
"description": "Result of writing a project export to disk.",
"properties": {
"saved_to": {
"title": "Saved To",
"type": "string"
},
"file_count": {
"title": "File Count",
"type": "integer"
}
},
"required": [
"saved_to",
"file_count"
],
"title": "ProjectSaveResponse",
"type": "object"
}
Fields:
-
saved_to(str) -
file_count(int)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
213 214 215 216 217 | |
RunInformation
pydantic-model
Bases: BaseModel
Contains summary information about a complete FlowGraph execution.
Show JSON schema:
{
"$defs": {
"NodeResult": {
"description": "Represents the execution result of a single node in a FlowGraph run.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"node_name": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Node Name"
},
"description": {
"default": "",
"title": "Description",
"type": "string"
},
"start_timestamp": {
"title": "Start Timestamp",
"type": "number"
},
"end_timestamp": {
"default": 0,
"title": "End Timestamp",
"type": "number"
},
"success": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Success"
},
"error": {
"default": "",
"title": "Error",
"type": "string"
},
"run_time_ms": {
"default": -1,
"description": "Run time in milliseconds",
"title": "Run Time Ms",
"type": "integer"
},
"is_running": {
"default": true,
"title": "Is Running",
"type": "boolean"
}
},
"required": [
"node_id"
],
"title": "NodeResult",
"type": "object"
}
},
"description": "Contains summary information about a complete FlowGraph execution.",
"properties": {
"flow_id": {
"title": "Flow Id",
"type": "integer"
},
"start_time": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"title": "Start Time"
},
"end_time": {
"anyOf": [
{
"format": "date-time",
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "End Time"
},
"success": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"title": "Success"
},
"is_running": {
"default": false,
"title": "Is Running",
"type": "boolean"
},
"execution_mode": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Execution Mode"
},
"nodes_completed": {
"default": 0,
"title": "Nodes Completed",
"type": "integer"
},
"number_of_nodes": {
"default": 0,
"title": "Number Of Nodes",
"type": "integer"
},
"node_step_result": {
"items": {
"$ref": "#/$defs/NodeResult"
},
"title": "Node Step Result",
"type": "array"
},
"run_type": {
"enum": [
"fetch_one",
"full_run",
"init"
],
"title": "Run Type",
"type": "string"
}
},
"required": [
"flow_id",
"node_step_result",
"run_type"
],
"title": "RunInformation",
"type": "object"
}
Fields:
-
flow_id(int) -
start_time(datetime | None) -
end_time(datetime | None) -
success(bool | None) -
is_running(bool) -
execution_mode(str | None) -
nodes_completed(int) -
number_of_nodes(int) -
node_step_result(list[NodeResult]) -
run_type(Literal['fetch_one', 'full_run', 'init'])
Source code in flowfile_core/flowfile_core/schemas/output_model.py
28 29 30 31 32 33 34 35 36 37 38 39 40 | |
TableExample
pydantic-model
Bases: BaseModel
Represents a preview of a table, including schema and sample data.
number_of_records is None when the total is unknown (e.g. a lazy result
whose count was never computed); 0 always means a genuinely empty result.
Show JSON schema:
{
"$defs": {
"FileColumn": {
"description": "Represents detailed schema and statistics for a single column (field).\n\nThe statistics fields are None until they are actually computed \u2014 either\nnever (plain schema previews) or exactly, on demand, via the column-stats\nendpoint writing into the node's ``FlowfileColumn``.",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"title": "Data Type",
"type": "string"
},
"data_type_group": {
"default": "Other",
"enum": [
"Numeric",
"String",
"Date",
"Other",
"Boolean",
"Binary",
"Complex"
],
"title": "Data Type Group",
"type": "string"
},
"is_unique": {
"default": false,
"title": "Is Unique",
"type": "boolean"
},
"max_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Max Value"
},
"min_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Min Value"
},
"average_value": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Average Value"
},
"number_of_empty_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Empty Values"
},
"number_of_filled_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Filled Values"
},
"number_of_unique_values": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Unique Values"
},
"size": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Size"
}
},
"required": [
"name",
"data_type"
],
"title": "FileColumn",
"type": "object"
}
},
"description": "Represents a preview of a table, including schema and sample data.\n\n``number_of_records`` is None when the total is unknown (e.g. a lazy result\nwhose count was never computed); 0 always means a genuinely empty result.",
"properties": {
"node_id": {
"title": "Node Id",
"type": "integer"
},
"number_of_records": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Number Of Records"
},
"number_of_columns": {
"title": "Number Of Columns",
"type": "integer"
},
"name": {
"title": "Name",
"type": "string"
},
"table_schema": {
"items": {
"$ref": "#/$defs/FileColumn"
},
"title": "Table Schema",
"type": "array"
},
"columns": {
"items": {
"type": "string"
},
"title": "Columns",
"type": "array"
},
"data": {
"anyOf": [
{
"items": {
"additionalProperties": true,
"type": "object"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Data"
},
"has_example_data": {
"default": false,
"title": "Has Example Data",
"type": "boolean"
},
"has_run_with_current_setup": {
"default": false,
"title": "Has Run With Current Setup",
"type": "boolean"
}
},
"required": [
"node_id",
"number_of_columns",
"name",
"table_schema",
"columns"
],
"title": "TableExample",
"type": "object"
}
Fields:
-
node_id(int) -
number_of_records(int | None) -
number_of_columns(int) -
name(str) -
table_schema(list[FileColumn]) -
columns(list[str]) -
data(list[dict] | None) -
has_example_data(bool) -
has_run_with_current_setup(bool)
Source code in flowfile_core/flowfile_core/schemas/output_model.py
77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 | |
Web API
This section documents the FastAPI routes that expose flowfile-core's functionality over HTTP.
routes
flowfile_core.routes.routes
Main API router and endpoint definitions for the Flowfile application.
This module sets up the FastAPI router, defines all the API endpoints for interacting with flows, nodes, files, and other core components of the application. It handles the logic for creating, reading, updating, and deleting these resources.
Classes:
| Name | Description |
|---|---|
DynamicRenamePreviewRequest |
Request body for |
DynamicRenamePreviewResponse |
Response body for |
GroupOperationResponse |
OperationResponse that also returns the affected group (for server-assigned ids). |
RestApiSampleResponse |
Inferred output columns from a REST API sample fetch. |
Functions:
| Name | Description |
|---|---|
add_generic_settings |
A generic endpoint to update the settings of any node. |
add_node |
Adds a new, unconfigured node (a "promise") to the flow graph. |
add_nodes_to_group |
Add nodes to an existing group. |
cancel_flow |
Cancels a currently running flow execution. |
check_flow_laziness |
Check whether a flow supports fully lazy execution for virtual tables. |
clear_history |
Clear all history for a flow. |
close_flow |
Closes an active flow session for the current user. Idempotent: closing a flow that isn't in |
compute_node_visualization |
Compute Graphic Walker chart rows for an Explore Data node. |
connect_node |
Creates a connection (edge) between two nodes in the flow graph. |
copy_node |
Copies an existing node's settings to a new node promise. |
create_db_connection |
Creates and securely stores a new database connection. |
create_directory |
Creates a new directory at the specified path. |
create_flow |
Creates a new, empty flow file at the specified path and registers a session for it. |
create_from_template |
Instantiates a template as a new flow session. |
create_group |
Create a visual group around a set of nodes. Returns the new server-assigned group. |
delete_db_connection |
Deletes a stored database connection (own, or group-shared with manage access). |
delete_group |
Delete a group box (ungroup). Member nodes are kept. |
delete_node |
Deletes a node from the flow graph. |
delete_node_connection |
Deletes a connection (edge) between two nodes. |
download_generated_project |
Generates the project export and returns it as a zip archive. |
ensure_templates_available |
Downloads template flow YAMLs from GitHub if not already cached locally. |
fetch_rest_api_sample |
Fetch a small sample from the configured REST API and infer its schema. |
get_active_flow_file_sessions |
Retrieves a list of all currently active flow sessions for the current user. |
get_catalog_flows_directory |
Returns the managed flows directory used for catalog-tab saves. |
get_db_connections |
Retrieves all stored database connections for the current user (without passwords). |
get_db_dialects |
Returns the supported database dialects (drives the frontend's dialect dropdowns). |
get_db_schemas |
Returns available schema names for the given database connection. |
get_db_tables |
Returns available table names for the given database connection and optional schema. |
get_default_path |
Returns the default starting path for the file browser (user data directory). |
get_description_node |
Retrieves the description text for a specific node. |
get_directory_contents |
Gets the contents of a directory path. |
get_downstream_node_ids |
Gets a list of all node IDs that are downstream dependencies of a given node. |
get_excel_sheet_names |
Retrieves the sheet names from an Excel file. |
get_expression_doc |
Retrieves documentation for available Polars expressions. |
get_expressions |
Retrieves a list of all available Flowfile expression names. |
get_flow |
Retrieves the settings for a specific flow (including runtime dirty state). |
get_flow_artifacts |
Returns artifact visualization data for the canvas. |
get_flow_frontend_data |
Retrieves the data needed to render the flow graph in the frontend. |
get_flow_settings |
Retrieves the main settings for a flow (including dirty-state info). |
get_flow_settings_validation |
Conservative static check: node settings that reference missing input columns. |
get_generated_code |
Generates and returns a Python script with Polars code representing the flow. |
get_generated_flowframe_code |
Generates and returns a Python script with FlowFrame code representing the flow. |
get_generated_project |
Generates a multi-file Python project (FlowFrame code) representing the flow. |
get_graphic_walker_input |
Gets the saved chart specs and field schema for the Graphic Walker explorer. |
get_history_status |
Get the current state of the history system for a flow. |
get_instant_function_result |
Executes a simple, instant function on a node's data and returns the result. |
get_list_of_saved_flows |
Scans a directory for saved flow files ( |
get_local_files |
Retrieves a list of files from a specified local directory. |
get_node |
Retrieves the complete state and data preview for a single node. |
get_node_available_artifacts |
Return available artifact metadata for a node. |
get_node_column_stats |
Computes on-demand statistics for one column of a node's cached result. |
get_node_input_names |
Returns the named inputs available for a kernel node. |
get_node_list |
Retrieves the list of all available node types and their templates. |
get_node_model |
(Internal) Retrieves a node's Pydantic model from the input_schema module by its name. |
get_node_upstream_ids |
Return the transitive upstream node IDs for a given node. |
get_node_visualization_fields |
Return the Graphic Walker field schema for an Explore Data node's result. |
get_reference_node |
Retrieves the reference identifier for a specific node. |
get_run_status |
Retrieves the run status information for a specific flow. |
get_table_example |
Retrieves a data preview (schema and sample rows) for a node's output. |
get_vue_flow_data |
Retrieves the flow data formatted for the Vue-based frontend. |
import_saved_flow |
Imports a flow from a saved |
list_templates |
Returns metadata for all available flow templates. |
overwrite_flow_in_catalog |
Overwrite an existing catalog flow's YAML with the contents of another flow. |
preview_dynamic_rename |
Resolves a dynamic-rename rule against a given schema without mutating any flow. |
redo_action |
Redo the last undone action on the flow graph. |
register_flow |
Registers a new flow session with the application for the current user. |
remove_nodes_from_group |
Remove nodes from their group; a group emptied this way is pruned. |
rename_flow |
Renames a flow's display name: the catalog registration (when one exists) plus the |
run_flow |
Executes a flow in a background task. |
save_flow |
Deprecated GET variant of |
save_flow_post |
Saves the current state of a flow to a |
save_flow_to_catalog |
Save a flow into the managed catalog flows directory with a collision-free filename. |
save_generated_project |
Generates the project export and writes it into a directory on the server. |
trigger_fetch_node_data |
Fetches and refreshes the data for a specific node. |
undo_action |
Undo the last action on the flow graph. |
update_db_connection |
Updates an existing database connection (own, or group-shared with manage access). |
update_description_node |
Updates the description text for a specific node. |
update_flow_settings |
Updates the main settings for a flow. |
update_group |
Rename / recolor / move / resize / collapse a group box. |
update_layout |
Persist dragged node positions and/or group bounds (one drag-end -> one call). |
update_reference_node |
Updates the reference identifier for a specific node. |
validate_db_settings |
Validates that a connection can be made to a database with the given settings. |
validate_node_reference |
Validates if a reference is valid and unique for a node. |
DynamicRenamePreviewRequest
pydantic-model
Bases: BaseModel
Request body for /dynamic_rename/preview.
Show JSON schema:
{
"$defs": {
"DynamicRenameInput": {
"description": "Defines settings for a dynamic rename operation.\n\nApplies a single rule (prefix / suffix / formula / first_row) to a set of selected\ncolumns, rather than requiring the user to rename columns one-by-one.\n\nIn formula mode, the flowfile formula syntax is evaluated with `[column_name]`\nbound to each target column's current name; for example `uppercase([column_name])`\nor `\"v2_\" + [column_name]`.\n\nIn first_row mode, the first row of the incoming table is promoted to column\nheaders and then dropped from the data. Non-string values are coerced to `str`;\nnull or empty values raise an error. Selection filters still apply \u2014 only selected\ncolumns are renamed, but the first row is always dropped.",
"properties": {
"rename_mode": {
"default": "prefix",
"enum": [
"prefix",
"suffix",
"formula",
"first_row"
],
"title": "Rename Mode",
"type": "string"
},
"prefix": {
"default": "",
"title": "Prefix",
"type": "string"
},
"suffix": {
"default": "",
"title": "Suffix",
"type": "string"
},
"formula": {
"default": "",
"expression": true,
"title": "Formula",
"type": "string"
},
"selection_mode": {
"default": "all",
"enum": [
"all",
"list",
"data_type"
],
"title": "Selection Mode",
"type": "string"
},
"selected_columns": {
"items": {
"type": "string"
},
"title": "Selected Columns",
"type": "array"
},
"selected_data_type": {
"anyOf": [
{
"enum": [
"Numeric",
"String",
"Date",
"Other",
"Boolean",
"Binary",
"Complex"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Selected Data Type"
}
},
"title": "DynamicRenameInput",
"type": "object"
},
"_DynamicRenameColumn": {
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type_group": {
"default": "",
"title": "Data Type Group",
"type": "string"
}
},
"required": [
"name"
],
"title": "_DynamicRenameColumn",
"type": "object"
}
},
"description": "Request body for `/dynamic_rename/preview`.",
"properties": {
"settings": {
"$ref": "#/$defs/DynamicRenameInput"
},
"incoming_columns": {
"items": {
"$ref": "#/$defs/_DynamicRenameColumn"
},
"title": "Incoming Columns",
"type": "array"
}
},
"required": [
"settings"
],
"title": "DynamicRenamePreviewRequest",
"type": "object"
}
Fields:
-
settings(DynamicRenameInput) -
incoming_columns(list[_DynamicRenameColumn])
Source code in flowfile_core/flowfile_core/routes/routes.py
1559 1560 1561 1562 1563 | |
DynamicRenamePreviewResponse
pydantic-model
Bases: BaseModel
Response body for /dynamic_rename/preview.
Show JSON schema:
{
"description": "Response body for `/dynamic_rename/preview`.",
"properties": {
"rename_map": {
"additionalProperties": {
"type": "string"
},
"title": "Rename Map",
"type": "object"
},
"error": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Error"
}
},
"required": [
"rename_map"
],
"title": "DynamicRenamePreviewResponse",
"type": "object"
}
Fields:
-
rename_map(dict[str, str]) -
error(str | None)
Source code in flowfile_core/flowfile_core/routes/routes.py
1566 1567 1568 1569 1570 | |
GroupOperationResponse
pydantic-model
Bases: OperationResponse
OperationResponse that also returns the affected group (for server-assigned ids).
Show JSON schema:
{
"$defs": {
"FlowfileGroup": {
"description": "Serialized representation of a visual node group (YAML/JSON).",
"properties": {
"id": {
"title": "Id",
"type": "integer"
},
"name": {
"default": "Group",
"title": "Name",
"type": "string"
},
"color": {
"anyOf": [
{
"enum": [
"slate",
"blue",
"green",
"amber",
"rose",
"violet",
"cyan"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Color"
},
"x_position": {
"default": 0.0,
"title": "X Position",
"type": "number"
},
"y_position": {
"default": 0.0,
"title": "Y Position",
"type": "number"
},
"width": {
"default": 400.0,
"title": "Width",
"type": "number"
},
"height": {
"default": 250.0,
"title": "Height",
"type": "number"
},
"collapsed": {
"default": false,
"title": "Collapsed",
"type": "boolean"
},
"parent_group_id": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"title": "Parent Group Id"
}
},
"required": [
"id"
],
"title": "FlowfileGroup",
"type": "object"
},
"HistoryState": {
"description": "Current state of the history system.\n\nProvides information about what undo/redo operations are available.",
"properties": {
"can_undo": {
"default": false,
"description": "Whether undo is available",
"title": "Can Undo",
"type": "boolean"
},
"can_redo": {
"default": false,
"description": "Whether redo is available",
"title": "Can Redo",
"type": "boolean"
},
"undo_description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Description of the action that would be undone",
"title": "Undo Description"
},
"redo_description": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Description of the action that would be redone",
"title": "Redo Description"
},
"undo_count": {
"default": 0,
"description": "Number of available undo steps",
"title": "Undo Count",
"type": "integer"
},
"redo_count": {
"default": 0,
"description": "Number of available redo steps",
"title": "Redo Count",
"type": "integer"
}
},
"title": "HistoryState",
"type": "object"
}
},
"description": "OperationResponse that also returns the affected group (for server-assigned ids).",
"properties": {
"success": {
"default": true,
"description": "Whether the operation succeeded",
"title": "Success",
"type": "boolean"
},
"message": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Optional message",
"title": "Message"
},
"history": {
"$ref": "#/$defs/HistoryState",
"description": "Current history state after the operation"
},
"group": {
"anyOf": [
{
"$ref": "#/$defs/FlowfileGroup"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"history"
],
"title": "GroupOperationResponse",
"type": "object"
}
Fields:
-
success(bool) -
message(str | None) -
history(HistoryState) -
group(FlowfileGroup | None)
Source code in flowfile_core/flowfile_core/routes/routes.py
855 856 857 858 | |
history
pydantic-field
Current history state after the operation
message = None
pydantic-field
Optional message
success = True
pydantic-field
Whether the operation succeeded
RestApiSampleResponse
pydantic-model
Bases: BaseModel
Inferred output columns from a REST API sample fetch.
Show JSON schema:
{
"$defs": {
"MinimalFieldInfo": {
"description": "Represents the most basic information about a data field (column).",
"properties": {
"name": {
"title": "Name",
"type": "string"
},
"data_type": {
"default": "String",
"title": "Data Type",
"type": "string"
}
},
"required": [
"name"
],
"title": "MinimalFieldInfo",
"type": "object"
}
},
"description": "Inferred output columns from a REST API sample fetch.",
"properties": {
"fields": {
"items": {
"$ref": "#/$defs/MinimalFieldInfo"
},
"title": "Fields",
"type": "array"
}
},
"required": [
"fields"
],
"title": "RestApiSampleResponse",
"type": "object"
}
Fields:
-
fields(list[MinimalFieldInfo])
Source code in flowfile_core/flowfile_core/routes/routes.py
1467 1468 1469 1470 | |
add_generic_settings(input_data, node_type, current_user=Depends(get_current_active_user))
A generic endpoint to update the settings of any node.
This endpoint dynamically determines the correct Pydantic model and update
function based on the node_type parameter.
Returns:
| Type | Description |
|---|---|
OperationResponse
|
OperationResponse with current history state. |
Source code in flowfile_core/flowfile_core/routes/routes.py
1415 1416 1417 1418 1419 1420 1421 1422 1423 1424 1425 1426 1427 1428 1429 1430 1431 1432 1433 1434 1435 1436 1437 1438 1439 1440 1441 1442 1443 1444 1445 1446 1447 1448 1449 1450 1451 1452 1453 1454 1455 1456 1457 1458 1459 1460 1461 1462 1463 1464 | |
add_node(flow_id, node_id, node_type, pos_x=0, pos_y=0)
Adds a new, unconfigured node (a "promise") to the flow graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the flow to add the node to. |
required |
node_id
|
int
|
The client-generated ID for the new node. |
required |
node_type
|
str
|
The type of the node to add (e.g., 'filter', 'join'). |
required |
pos_x
|
int | float
|
The X coordinate for the node's position in the UI. |
0
|
pos_y
|
int | float
|
The Y coordinate for the node's position in the UI. |
0
|
Returns:
| Type | Description |
|---|---|
OperationResponse | None
|
OperationResponse with current history state. |
Source code in flowfile_core/flowfile_core/routes/routes.py
609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 | |
add_nodes_to_group(flow_id, group_id, request)
Add nodes to an existing group.
Source code in flowfile_core/flowfile_core/routes/routes.py
922 923 924 925 926 927 928 929 930 | |
cancel_flow(flow_id)
Cancels a currently running flow execution.
Source code in flowfile_core/flowfile_core/routes/routes.py
508 509 510 511 512 513 514 | |
check_flow_laziness(flow_id)
Check whether a flow supports fully lazy execution for virtual tables.
Source code in flowfile_core/flowfile_core/routes/routes.py
976 977 978 979 980 981 982 983 | |
clear_history(flow_id)
Clear all history for a flow.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the flow to clear history for. |
required |
Source code in flowfile_core/flowfile_core/routes/routes.py
1302 1303 1304 1305 1306 1307 1308 1309 1310 1311 1312 1313 | |
close_flow(flow_id, current_user=Depends(get_current_active_user))
Closes an active flow session for the current user. Idempotent: closing a flow that isn't in the session (e.g. a stale tab after a restore pruned it) is a no-op, not a 500.
Source code in flowfile_core/flowfile_core/routes/routes.py
1168 1169 1170 1171 1172 1173 1174 1175 | |
compute_node_visualization(body, current_user=Depends(get_current_active_user))
Compute Graphic Walker chart rows for an Explore Data node.
GW's computation callback posts its IDataQueryPayload here on every
aggregation; the worker's session cache keeps the node's lazy frame warm so
successive calls skip the load.
Source code in flowfile_core/flowfile_core/routes/routes.py
2366 2367 2368 2369 2370 2371 2372 2373 2374 2375 2376 2377 2378 2379 2380 2381 2382 2383 2384 2385 | |
connect_node(flow_id, node_connection)
Creates a connection (edge) between two nodes in the flow graph.
Returns:
| Type | Description |
|---|---|
OperationResponse
|
OperationResponse with current history state. |
Source code in flowfile_core/flowfile_core/routes/routes.py
829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 | |
copy_node(node_id_to_copy_from, flow_id_to_copy_from, node_promise)
Copies an existing node's settings to a new node promise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
node_id_to_copy_from
|
int
|
The ID of the node to copy the settings from. |
required |
flow_id_to_copy_from
|
int
|
The ID of the flow containing the source node. |
required |
node_promise
|
NodePromise
|
A |
required |
Returns:
| Type | Description |
|---|---|
OperationResponse
|
OperationResponse with current history state. |
Source code in flowfile_core/flowfile_core/routes/routes.py
562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 | |
create_db_connection(input_connection, current_user=Depends(get_current_active_user), db=Depends(get_db))
Creates and securely stores a new database connection.
Source code in flowfile_core/flowfile_core/routes/routes.py
751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 | |
create_directory(new_directory)
Creates a new directory at the specified path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
new_directory
|
NewDirectory
|
An |
required |
Returns:
| Type | Description |
|---|---|
bool
|
|
Source code in flowfile_core/flowfile_core/routes/routes.py
213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 | |
create_flow(flow_path=None, name=None, namespace_id=None, register_in_catalog=True, persist=True, current_user=Depends(get_current_active_user))
Creates a new, empty flow file at the specified path and registers a session for it.
Two independent switches, deliberately not fused:
persistwrites the YAML to disk.Falsekeeps the flow in-memory until an explicit save or first run, so an abandoned blank canvas leaves no orphan file.register_in_catalogcreates theFlowRegistrationrow.Falseyields a flow with nosource_registration_id: no run history, schedules or API publishing until it is filed. Whennamespace_idis provided the flow lands there; otherwise it auto-registers underGeneral > {Unnamed | Local} Flows.
Source code in flowfile_core/flowfile_core/routes/routes.py
1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 1108 1109 1110 1111 1112 1113 1114 1115 1116 1117 1118 1119 1120 1121 1122 1123 1124 1125 1126 1127 1128 1129 1130 1131 1132 1133 1134 1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 | |
create_from_template(template_id, current_user=Depends(get_current_active_user))
Instantiates a template as a new flow session.
Downloads required CSV data files from GitHub if not already cached locally, then creates a flow from the template definition.
The new flow is not registered in the catalog: the only path available to register is the temp file this route unlinks on the way out, so doing so minted a permanently dangling row. Saving the flow files it.
Source code in flowfile_core/flowfile_core/routes/routes.py
2495 2496 2497 2498 2499 2500 2501 2502 2503 2504 2505 2506 2507 2508 2509 2510 2511 2512 2513 2514 2515 2516 2517 2518 2519 2520 2521 2522 2523 2524 2525 2526 2527 2528 2529 2530 2531 2532 2533 2534 2535 2536 2537 2538 2539 2540 2541 2542 2543 2544 2545 2546 2547 2548 2549 2550 2551 2552 2553 2554 2555 2556 2557 2558 2559 2560 2561 2562 2563 2564 2565 2566 2567 2568 2569 2570 2571 2572 2573 2574 2575 2576 2577 2578 | |
create_group(flow_id, request)
Create a visual group around a set of nodes. Returns the new server-assigned group.
Source code in flowfile_core/flowfile_core/routes/routes.py
882 883 884 885 886 887 888 889 890 891 892 893 894 | |
delete_db_connection(connection_name, current_user=Depends(get_current_active_user), db=Depends(get_db))
Deletes a stored database connection (own, or group-shared with manage access).
Source code in flowfile_core/flowfile_core/routes/routes.py
805 806 807 808 809 810 811 812 813 814 815 816 | |
delete_group(flow_id, group_id)
Delete a group box (ungroup). Member nodes are kept.
Source code in flowfile_core/flowfile_core/routes/routes.py
914 915 916 917 918 919 | |
delete_node(flow_id, node_id)
Deletes a node from the flow graph.
Returns:
| Type | Description |
|---|---|
OperationResponse
|
OperationResponse with current history state. |
Source code in flowfile_core/flowfile_core/routes/routes.py
691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 | |
delete_node_connection(flow_id, node_connection=None)
Deletes a connection (edge) between two nodes.
Returns:
| Type | Description |
|---|---|
OperationResponse
|
OperationResponse with current history state. |
Source code in flowfile_core/flowfile_core/routes/routes.py
712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 | |
download_generated_project(flow_id)
Generates the project export and returns it as a zip archive.
Source code in flowfile_core/flowfile_core/routes/routes.py
1029 1030 1031 1032 1033 1034 1035 1036 1037 | |
ensure_templates_available()
Downloads template flow YAMLs from GitHub if not already cached locally.
Called by the frontend on first visit to the templates page to ensure templates are available even when running from a PyPI install (no repo checkout).
Source code in flowfile_core/flowfile_core/routes/routes.py
2478 2479 2480 2481 2482 2483 2484 2485 2486 2487 2488 2489 2490 2491 2492 | |
fetch_rest_api_sample(input_data, sample_size=50, current_user=Depends(get_current_active_user))
Fetch a small sample from the configured REST API and infer its schema.
Runs one capped request through the worker (so all network I/O stays
sandboxed off the core event loop), infers the output columns with Polars,
caches them on the node's fields (so downstream schema prediction needs
no network), and returns the inferred columns. This powers the node's
"Fetch sample" button. Defined as a sync endpoint so the blocking worker
round-trip runs in FastAPI's threadpool rather than the event loop.
Source code in flowfile_core/flowfile_core/routes/routes.py
1473 1474 1475 1476 1477 1478 1479 1480 1481 1482 1483 1484 1485 1486 1487 1488 1489 1490 1491 1492 1493 1494 1495 1496 1497 1498 1499 1500 1501 1502 1503 1504 1505 1506 1507 1508 1509 1510 1511 1512 1513 1514 1515 1516 1517 | |
get_active_flow_file_sessions(current_user=Depends(get_current_active_user))
async
Retrieves a list of all currently active flow sessions for the current user.
Source code in flowfile_core/flowfile_core/routes/routes.py
244 245 246 247 248 249 250 251 252 253 | |
get_catalog_flows_directory()
async
Returns the managed flows directory used for catalog-tab saves.
On local this resolves to ~/.flowfile/flows; in Docker mode to
/data/user/flows. The frontend uses this to build the target path
for flows saved via the Catalog tab, so they always land in the managed
location regardless of where the file browser was last navigated.
Source code in flowfile_core/flowfile_core/routes/routes.py
173 174 175 176 177 178 179 180 181 182 | |
get_db_connections(db=Depends(get_db), current_user=Depends(get_current_active_user))
Retrieves all stored database connections for the current user (without passwords).
Source code in flowfile_core/flowfile_core/routes/routes.py
819 820 821 822 823 824 825 826 | |
get_db_dialects()
Returns the supported database dialects (drives the frontend's dialect dropdowns).
Source code in flowfile_core/flowfile_core/routes/routes.py
737 738 739 740 | |
get_db_schemas(database_settings, current_user=Depends(get_current_active_user))
async
Returns available schema names for the given database connection.
Source code in flowfile_core/flowfile_core/routes/routes.py
2440 2441 2442 2443 2444 2445 2446 2447 2448 | |
get_db_tables(database_settings, current_user=Depends(get_current_active_user))
async
Returns available table names for the given database connection and optional schema.
When schema_name is provided, returns plain table names. When schema_name is not provided, returns schema-qualified names (schema.table) across all accessible schemas (skipping any that the user cannot read).
Source code in flowfile_core/flowfile_core/routes/routes.py
2451 2452 2453 2454 2455 2456 2457 2458 2459 2460 2461 2462 2463 2464 | |
get_default_path()
async
Returns the default starting path for the file browser (user data directory).
Source code in flowfile_core/flowfile_core/routes/routes.py
167 168 169 170 | |
get_description_node(flow_id, node_id)
Retrieves the description text for a specific node.
Returns the user-provided description if set, otherwise falls back
to an auto-generated description based on the node's configuration.
The response includes an is_auto_generated flag so the frontend
knows whether to refresh the description after settings changes.
Source code in flowfile_core/flowfile_core/routes/routes.py
1669 1670 1671 1672 1673 1674 1675 1676 1677 1678 1679 1680 1681 1682 1683 1684 1685 1686 1687 1688 1689 1690 | |
get_directory_contents(directory, file_types=None, include_hidden=False)
async
Gets the contents of a directory path.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
directory
|
str
|
The absolute path to the directory. |
required |
file_types
|
list[str]
|
An optional list of file extensions to filter by. |
None
|
include_hidden
|
bool
|
If True, includes hidden files and directories. |
False
|
Returns:
| Type | Description |
|---|---|
list[FileInfo]
|
A list of |
Source code in flowfile_core/flowfile_core/routes/routes.py
185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 | |
get_downstream_node_ids(flow_id, node_id)
async
Gets a list of all node IDs that are downstream dependencies of a given node.
Source code in flowfile_core/flowfile_core/routes/routes.py
1818 1819 1820 1821 1822 1823 | |
get_excel_sheet_names(path)
async
Retrieves the sheet names from an Excel file.
Source code in flowfile_core/flowfile_core/routes/routes.py
2416 2417 2418 2419 2420 2421 2422 2423 2424 | |
get_expression_doc()
Retrieves documentation for available Polars expressions.
Source code in flowfile_core/flowfile_core/routes/routes.py
956 957 958 959 | |
get_expressions()
Retrieves a list of all available Flowfile expression names.
Source code in flowfile_core/flowfile_core/routes/routes.py
962 963 964 965 | |
get_flow(flow_id)
Retrieves the settings for a specific flow (including runtime dirty state).
Source code in flowfile_core/flowfile_core/routes/routes.py
968 969 970 971 972 973 | |
get_flow_artifacts(flow_id)
Returns artifact visualization data for the canvas.
Includes per-node artifact summaries (for badges/tooltips) and artifact edges (for dashed-line connections between publisher and consumer nodes).
Source code in flowfile_core/flowfile_core/routes/routes.py
2237 2238 2239 2240 2241 2242 2243 2244 2245 2246 2247 2248 2249 2250 2251 2252 | |
get_flow_frontend_data(flow_id=1)
Retrieves the data needed to render the flow graph in the frontend.
Source code in flowfile_core/flowfile_core/routes/routes.py
2200 2201 2202 2203 2204 2205 2206 | |
get_flow_settings(flow_id=1)
Retrieves the main settings for a flow (including dirty-state info).
Source code in flowfile_core/flowfile_core/routes/routes.py
2209 2210 2211 2212 2213 2214 2215 | |
get_flow_settings_validation(flow_id)
Conservative static check: node settings that reference missing input columns.
Source code in flowfile_core/flowfile_core/routes/routes.py
2255 2256 2257 2258 2259 2260 2261 | |
get_generated_code(flow_id)
Generates and returns a Python script with Polars code representing the flow.
Source code in flowfile_core/flowfile_core/routes/routes.py
986 987 988 989 990 991 992 993 994 995 996 | |
get_generated_flowframe_code(flow_id)
Generates and returns a Python script with FlowFrame code representing the flow.
Source code in flowfile_core/flowfile_core/routes/routes.py
999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 | |
get_generated_project(flow_id)
Generates a multi-file Python project (FlowFrame code) representing the flow.
Source code in flowfile_core/flowfile_core/routes/routes.py
1023 1024 1025 1026 | |
get_graphic_walker_input(flow_id, node_id, current_user=Depends(get_current_active_user))
Gets the saved chart specs and field schema for the Graphic Walker explorer.
Carries no rows: aggregation runs on the worker via /analysis_data/compute.
Source code in flowfile_core/flowfile_core/routes/routes.py
2336 2337 2338 2339 2340 2341 2342 2343 2344 2345 2346 | |
get_history_status(flow_id)
Get the current state of the history system for a flow.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the flow to get history status for. |
required |
Returns:
| Type | Description |
|---|---|
HistoryState
|
HistoryState with information about available undo/redo operations. |
Source code in flowfile_core/flowfile_core/routes/routes.py
1286 1287 1288 1289 1290 1291 1292 1293 1294 1295 1296 1297 1298 1299 | |
get_instant_function_result(flow_id, node_id, func_string)
async
Executes a simple, instant function on a node's data and returns the result.
Source code in flowfile_core/flowfile_core/routes/routes.py
2405 2406 2407 2408 2409 2410 2411 2412 2413 | |
get_list_of_saved_flows(path)
Scans a directory for saved flow files (.flowfile).
Source code in flowfile_core/flowfile_core/routes/routes.py
1520 1521 1522 1523 1524 1525 1526 1527 1528 1529 1530 1531 1532 1533 | |
get_local_files(directory)
async
Retrieves a list of files from a specified local directory.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
directory
|
str
|
The absolute path of the directory to scan. |
required |
Returns:
| Type | Description |
|---|---|
list[FileInfo]
|
A list of |
Raises:
| Type | Description |
|---|---|
HTTPException
|
404 if the directory does not exist. |
HTTPException
|
403 if access is denied (path outside sandbox). |
Source code in flowfile_core/flowfile_core/routes/routes.py
141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 | |
get_node(flow_id, node_id, get_data=False, include_output=True, include_inputs=True)
Retrieves the complete state and data preview for a single node.
When include_output is False the node's own output preview
(main_output) is skipped. The settings panel only needs the input
schemas, and computing the output can be expensive for data-dependent nodes
(e.g. a pivot must materialize data to determine its output columns), so the
editor opens settings with include_output=false for an instant response.
When include_inputs is also False the input schemas are skipped too:
resolving them can execute un-run upstream custom nodes (kernel/worker).
The custom-node drawer uses this to render settings instantly and hydrates
columns with a follow-up full fetch.
Source code in flowfile_core/flowfile_core/routes/routes.py
1598 1599 1600 1601 1602 1603 1604 1605 1606 1607 1608 1609 1610 1611 1612 1613 1614 1615 1616 1617 1618 1619 1620 1621 1622 1623 1624 1625 1626 | |
get_node_available_artifacts(flow_id, node_id, kernel_id=None)
Return available artifact metadata for a node.
Merges run-observed artifacts (published in a prior run) with artifacts upstream kernel nodes declare they publish (from their manifests), so the frontend's artifact pickers work before the flow has ever run. Observed wins on name conflict.
Source code in flowfile_core/flowfile_core/routes/routes.py
2291 2292 2293 2294 2295 2296 2297 2298 2299 2300 2301 2302 2303 2304 2305 2306 2307 2308 2309 2310 2311 2312 2313 2314 2315 2316 2317 2318 2319 2320 2321 2322 2323 2324 2325 2326 2327 2328 2329 2330 2331 2332 2333 | |
get_node_column_stats(flow_id, node_id, column_name, output_handle=DEFAULT_OUTPUT_HANDLE)
Computes on-demand statistics for one column of a node's cached result.
Runs a single bounded aggregate (counts, uniques, min/max) over the result
the last run left behind — it never re-executes the node — and returns the
column's updated FileColumn, the same shape table_schema ships.
Query parameters (not path segments): column names contain / and ..
Source code in flowfile_core/flowfile_core/routes/routes.py
1785 1786 1787 1788 1789 1790 1791 1792 1793 1794 1795 1796 1797 1798 1799 1800 1801 1802 1803 1804 1805 1806 1807 1808 1809 1810 1811 1812 1813 1814 1815 | |
get_node_input_names(flow_id, node_id)
Returns the named inputs available for a kernel node.
Each entry contains the input name (derived from the source node's
node_reference or fallback df_{id}), the source node ID, and
its type. The frontend uses this for autocomplete and display.
Source code in flowfile_core/flowfile_core/routes/routes.py
1629 1630 1631 1632 1633 1634 1635 1636 1637 1638 1639 1640 1641 1642 1643 1644 1645 1646 1647 1648 1649 1650 1651 1652 1653 1654 1655 | |
get_node_list()
Retrieves the list of all available node types and their templates.
Source code in flowfile_core/flowfile_core/routes/routes.py
1536 1537 1538 1539 1540 1541 1542 1543 1544 1545 1546 1547 1548 1549 1550 1551 | |
get_node_model(setting_name_ref)
(Internal) Retrieves a node's Pydantic model from the input_schema module by its name.
Source code in flowfile_core/flowfile_core/routes/routes.py
131 132 133 134 135 136 137 138 | |
get_node_upstream_ids(flow_id, node_id)
Return the transitive upstream node IDs for a given node.
Used by the frontend to determine which artifacts are actually reachable (via the DAG) from a specific python_script node.
Source code in flowfile_core/flowfile_core/routes/routes.py
2264 2265 2266 2267 2268 2269 2270 2271 2272 2273 2274 | |
get_node_visualization_fields(body, current_user=Depends(get_current_active_user))
Return the Graphic Walker field schema for an Explore Data node's result.
Source code in flowfile_core/flowfile_core/routes/routes.py
2388 2389 2390 2391 2392 2393 2394 2395 2396 2397 2398 2399 2400 2401 2402 | |
get_reference_node(flow_id, node_id)
Retrieves the reference identifier for a specific node.
Source code in flowfile_core/flowfile_core/routes/routes.py
1730 1731 1732 1733 1734 1735 1736 1737 1738 1739 | |
get_run_status(flow_id, response)
Retrieves the run status information for a specific flow.
Returns a 202 Accepted status while the flow is running, and 200 OK when finished.
Source code in flowfile_core/flowfile_core/routes/routes.py
530 531 532 533 534 535 536 537 538 539 540 541 542 543 | |
get_table_example(flow_id, node_id, output_handle=DEFAULT_OUTPUT_HANDLE)
Retrieves a data preview (schema and sample rows) for a node's output.
For multi-output nodes, output_handle selects which named output to
preview (e.g. "output-0", "output-1"); the default is the first.
Source code in flowfile_core/flowfile_core/routes/routes.py
1773 1774 1775 1776 1777 1778 1779 1780 1781 1782 | |
get_vue_flow_data(flow_id)
Retrieves the flow data formatted for the Vue-based frontend.
Source code in flowfile_core/flowfile_core/routes/routes.py
2227 2228 2229 2230 2231 2232 2233 2234 | |
import_saved_flow(flow_path, current_user=Depends(get_current_active_user))
Imports a flow from a saved .yaml and registers it as a new session for the current user.
Opening a file is browsing, not filing: an existing registration is adopted so the
session gets its display name and source_registration_id, but no new catalog
row is created. Use the Save dialog (or POST /catalog/flows) to file a flow.
Source code in flowfile_core/flowfile_core/routes/routes.py
1826 1827 1828 1829 1830 1831 1832 1833 1834 1835 1836 1837 1838 1839 1840 1841 1842 | |
list_templates()
Returns metadata for all available flow templates.
Source code in flowfile_core/flowfile_core/routes/routes.py
2470 2471 2472 2473 2474 2475 | |
overwrite_flow_in_catalog(flow_id, target_registration_id, current_user=Depends(get_current_active_user))
Overwrite an existing catalog flow's YAML with the contents of another flow.
Unlike /save_flow_to_catalog, this intentionally writes over an existing
registration. The target registration's name and namespace are preserved;
only the file contents on disk change. Primary use case: reverting a flow
to an older version by loading that version and overwriting the canonical
catalog entry.
Returns the (possibly new) flow id so the frontend can switch to the target.
Source code in flowfile_core/flowfile_core/routes/routes.py
2114 2115 2116 2117 2118 2119 2120 2121 2122 2123 2124 2125 2126 2127 2128 2129 2130 2131 2132 2133 2134 2135 2136 2137 2138 2139 2140 2141 2142 2143 2144 2145 2146 2147 2148 2149 2150 2151 2152 2153 2154 2155 2156 2157 2158 2159 2160 2161 2162 2163 2164 2165 2166 2167 2168 2169 2170 2171 2172 2173 2174 2175 2176 2177 2178 2179 2180 2181 2182 2183 | |
preview_dynamic_rename(request)
Resolves a dynamic-rename rule against a given schema without mutating any flow.
The frontend calls this to render the live old-to-new preview pane inside the
node's settings panel. Returns either the fully-resolved rename map (possibly
empty) or an error describing a parse failure or duplicate-name collision.
first_row mode is intentionally not previewed here: its new names depend on
actual row data, and we don't want to trigger upstream computation from a
settings panel. The frontend renders a runtime-only placeholder for that mode.
Source code in flowfile_core/flowfile_core/routes/routes.py
1573 1574 1575 1576 1577 1578 1579 1580 1581 1582 1583 1584 1585 1586 1587 1588 1589 1590 1591 1592 1593 1594 1595 | |
redo_action(flow_id)
Redo the last undone action on the flow graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the flow to redo. |
required |
Returns:
| Type | Description |
|---|---|
UndoRedoResult
|
UndoRedoResult indicating success or failure. |
Source code in flowfile_core/flowfile_core/routes/routes.py
1268 1269 1270 1271 1272 1273 1274 1275 1276 1277 1278 1279 1280 1281 1282 1283 | |
register_flow(flow_data, current_user=Depends(get_current_active_user))
Registers a new flow session with the application for the current user.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_data
|
FlowSettings
|
The |
required |
Returns:
| Type | Description |
|---|---|
int
|
The ID of the newly registered flow. |
Source code in flowfile_core/flowfile_core/routes/routes.py
230 231 232 233 234 235 236 237 238 239 240 241 | |
remove_nodes_from_group(flow_id, request)
Remove nodes from their group; a group emptied this way is pruned.
Source code in flowfile_core/flowfile_core/routes/routes.py
933 934 935 936 937 938 | |
rename_flow(body, current_user=Depends(get_current_active_user))
Renames a flow's display name: the catalog registration (when one exists) plus the in-memory session name. The file path is never touched.
For an unregistered flow the name is written back into its YAML, because there is no
registration to hold it and open_flow would otherwise have nothing to read: the
rename would silently vanish when the tab is closed.
Source code in flowfile_core/flowfile_core/routes/routes.py
1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 1210 1211 1212 1213 1214 1215 1216 1217 1218 1219 1220 1221 1222 1223 1224 1225 1226 1227 1228 1229 1230 1231 1232 1233 1234 1235 1236 1237 1238 1239 1240 1241 1242 1243 1244 | |
run_flow(flow_id, background_tasks, current_user=Depends(get_current_active_user))
async
Executes a flow in a background task.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the flow to execute. |
required |
background_tasks
|
BackgroundTasks
|
FastAPI's background task runner. |
required |
Returns:
| Type | Description |
|---|---|
JSONResponse
|
A JSON response indicating that the flow has started. |
Source code in flowfile_core/flowfile_core/routes/routes.py
476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 | |
save_flow(response, flow_id, flow_path=None, namespace_id=None, register_in_catalog=True, current_user=Depends(get_current_active_user))
Deprecated GET variant of /save_flow. Prefer POST.
Kept for backward compatibility with older frontends/clients. Emits a
Deprecation: true response header.
Source code in flowfile_core/flowfile_core/routes/routes.py
1943 1944 1945 1946 1947 1948 1949 1950 1951 1952 1953 1954 1955 1956 1957 1958 1959 | |
save_flow_post(flow_id, flow_path=None, namespace_id=None, register_in_catalog=True, current_user=Depends(get_current_active_user))
Saves the current state of a flow to a .yaml.
See :func:_save_flow_impl for semantics.
Source code in flowfile_core/flowfile_core/routes/routes.py
1962 1963 1964 1965 1966 1967 1968 1969 1970 1971 1972 1973 1974 | |
save_flow_to_catalog(flow_id, flow_name, namespace_id, current_user=Depends(get_current_active_user))
Save a flow into the managed catalog flows directory with a collision-free filename.
The file is always written to {flows_dir}/{flow_id}_{sanitized_name}.yaml so
two flows with the same user-chosen name in different namespaces cannot overwrite
each other. Returns the (possibly new) flow id so the frontend can switch to it.
Source code in flowfile_core/flowfile_core/routes/routes.py
1977 1978 1979 1980 1981 1982 1983 1984 1985 1986 1987 1988 1989 1990 1991 1992 1993 1994 1995 1996 1997 1998 1999 2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2012 2013 2014 2015 2016 2017 2018 2019 2020 2021 2022 2023 2024 2025 2026 2027 2028 2029 2030 2031 2032 2033 2034 2035 2036 2037 2038 2039 2040 2041 2042 2043 2044 2045 2046 2047 2048 2049 2050 2051 2052 2053 2054 2055 2056 2057 2058 2059 2060 2061 2062 2063 2064 2065 2066 2067 2068 2069 2070 2071 2072 2073 2074 2075 2076 2077 2078 2079 2080 2081 2082 2083 2084 2085 2086 2087 2088 2089 2090 2091 2092 2093 2094 2095 2096 2097 2098 2099 2100 2101 2102 2103 2104 2105 2106 2107 2108 2109 2110 2111 | |
save_generated_project(request)
Generates the project export and writes it into a directory on the server.
The target directory is validated with the same sandbox rules as the file
browser (unrestricted in Electron mode, sandboxed to the user data
directory otherwise). Files are written under
<target_directory>/<project_name>/; existing project directories are
rejected with 409 unless overwrite is set. Existing files are only
overwritten file-by-file, never deleted.
Source code in flowfile_core/flowfile_core/routes/routes.py
1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 | |
trigger_fetch_node_data(flow_id, node_id, background_tasks, performance_mode=False)
async
Fetches and refreshes the data for a specific node.
performance_mode=true builds the node's query plan without storing its
result — enough for the Explore Data drawer, which charts through the worker
and never reads the example rows the default (preview) path materialises.
Source code in flowfile_core/flowfile_core/routes/routes.py
256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 | |
undo_action(flow_id)
Undo the last action on the flow graph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flow_id
|
int
|
The ID of the flow to undo. |
required |
Returns:
| Type | Description |
|---|---|
UndoRedoResult
|
UndoRedoResult indicating success or failure. |
Source code in flowfile_core/flowfile_core/routes/routes.py
1250 1251 1252 1253 1254 1255 1256 1257 1258 1259 1260 1261 1262 1263 1264 1265 | |
update_db_connection(input_connection, current_user=Depends(get_current_active_user), db=Depends(get_db))
Updates an existing database connection (own, or group-shared with manage access).
Source code in flowfile_core/flowfile_core/routes/routes.py
770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 | |
update_description_node(flow_id, node_id, description=Body(...))
Updates the description text for a specific node.
Source code in flowfile_core/flowfile_core/routes/routes.py
1658 1659 1660 1661 1662 1663 1664 1665 1666 | |
update_flow_settings(flow_settings)
Updates the main settings for a flow.
Source code in flowfile_core/flowfile_core/routes/routes.py
2218 2219 2220 2221 2222 2223 2224 | |
update_group(flow_id, group_id, request)
Rename / recolor / move / resize / collapse a group box.
Source code in flowfile_core/flowfile_core/routes/routes.py
897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 | |
update_layout(flow_id, request)
Persist dragged node positions and/or group bounds (one drag-end -> one call).
Also closes the long-standing gap where dragged node positions were never persisted.
Source code in flowfile_core/flowfile_core/routes/routes.py
941 942 943 944 945 946 947 948 949 950 951 952 953 | |
update_reference_node(flow_id, node_id, reference=Body(...))
Updates the reference identifier for a specific node.
The reference must be: - Lowercase only - No spaces allowed - Unique across all nodes in the flow
Source code in flowfile_core/flowfile_core/routes/routes.py
1693 1694 1695 1696 1697 1698 1699 1700 1701 1702 1703 1704 1705 1706 1707 1708 1709 1710 1711 1712 1713 1714 1715 1716 1717 1718 1719 1720 1721 1722 1723 1724 1725 1726 1727 | |
validate_db_settings(database_settings, current_user=Depends(get_current_active_user))
async
Validates that a connection can be made to a database with the given settings.
Source code in flowfile_core/flowfile_core/routes/routes.py
2427 2428 2429 2430 2431 2432 2433 2434 2435 2436 2437 | |
validate_node_reference(flow_id, node_id, reference)
Validates if a reference is valid and unique for a node.
Returns:
| Type | Description |
|---|---|
|
Dict with 'valid' (bool) and 'error' (str or None) fields. |
Source code in flowfile_core/flowfile_core/routes/routes.py
1742 1743 1744 1745 1746 1747 1748 1749 1750 1751 1752 1753 1754 1755 1756 1757 1758 1759 1760 1761 1762 1763 1764 1765 1766 1767 1768 1769 1770 | |
auth
flowfile_core.routes.auth
Functions:
| Name | Description |
|---|---|
change_own_password |
Change the current user's password |
create_user |
Create a new user (admin only) |
delete_user |
Delete a user (admin only) |
get_password_requirements |
Get password requirements for client-side validation |
list_users |
List all users (admin only) |
refresh_access_token |
Exchange a valid refresh token for a new access token and refresh token. |
update_user |
Update a user (admin only) |
change_own_password(password_data, current_user=Depends(get_current_active_user), db=Depends(get_db))
async
Change the current user's password
Source code in flowfile_core/flowfile_core/routes/auth.py
269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 | |
create_user(user_data, current_user=Depends(get_current_admin_user), db=Depends(get_db))
async
Create a new user (admin only)
Source code in flowfile_core/flowfile_core/routes/auth.py
111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 | |
delete_user(user_id, current_user=Depends(get_current_admin_user), db=Depends(get_db))
async
Delete a user (admin only)
Source code in flowfile_core/flowfile_core/routes/auth.py
221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 | |
get_password_requirements()
async
Get password requirements for client-side validation
Source code in flowfile_core/flowfile_core/routes/auth.py
301 302 303 304 | |
list_users(current_user=Depends(get_current_admin_user), db=Depends(get_db))
async
List all users (admin only)
Source code in flowfile_core/flowfile_core/routes/auth.py
93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 | |
refresh_access_token(refresh_token=Form(...), db=Depends(get_db))
async
Exchange a valid refresh token for a new access token and refresh token.
Source code in flowfile_core/flowfile_core/routes/auth.py
58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | |
update_user(user_id, user_data, current_user=Depends(get_current_admin_user), db=Depends(get_db))
async
Update a user (admin only)
Source code in flowfile_core/flowfile_core/routes/auth.py
158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 | |
cloud_connections
flowfile_core.routes.cloud_connections
Functions:
| Name | Description |
|---|---|
create_cloud_storage_connection |
Create a new cloud storage connection. |
delete_cloud_connection_with_connection_name |
Delete a cloud connection (own, or group-shared with manage access). |
get_cloud_connections |
Get all cloud storage connections for the current user. |
update_cloud_storage_connection |
Update an existing cloud storage connection (own, or group-shared with manage access). |
create_cloud_storage_connection(input_connection, current_user=Depends(get_current_active_user), db=Depends(get_db))
Create a new cloud storage connection. Parameters input_connection: FullCloudStorageConnection schema containing connection details current_user: User obtained from Depends(get_current_active_user) db: Session obtained from Depends(get_db) Returns Dict with a success message
Source code in flowfile_core/flowfile_core/routes/cloud_connections.py
29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 | |
delete_cloud_connection_with_connection_name(connection_name, current_user=Depends(get_current_active_user), db=Depends(get_db))
Delete a cloud connection (own, or group-shared with manage access).
Source code in flowfile_core/flowfile_core/routes/cloud_connections.py
103 104 105 106 107 108 109 110 111 112 113 114 115 116 | |
get_cloud_connections(db=Depends(get_db), current_user=Depends(get_current_active_user))
Get all cloud storage connections for the current user. Parameters db: Session obtained from Depends(get_db) current_user: User obtained from Depends(get_current_active_user)
Returns List[FullCloudStorageConnectionInterface]
Source code in flowfile_core/flowfile_core/routes/cloud_connections.py
119 120 121 122 123 124 125 126 127 128 129 130 131 132 | |
update_cloud_storage_connection(input_connection, current_user=Depends(get_current_active_user), db=Depends(get_db))
Update an existing cloud storage connection (own, or group-shared with manage access).
Source code in flowfile_core/flowfile_core/routes/cloud_connections.py
55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 | |
logs
flowfile_core.routes.logs
Functions:
| Name | Description |
|---|---|
add_log |
Adds a log message to the log file for a given flow_id. |
add_raw_log |
Adds a log message to the log file for a given flow_id. |
format_sse_message |
Format the data as a proper SSE message |
stream_logs |
Streams logs for a given flow_id using Server-Sent Events. |
add_log(flow_id, log_message)
async
Adds a log message to the log file for a given flow_id.
Source code in flowfile_core/flowfile_core/routes/logs.py
35 36 37 38 39 40 41 42 | |
add_raw_log(raw_log_input)
async
Adds a log message to the log file for a given flow_id.
Source code in flowfile_core/flowfile_core/routes/logs.py
45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 | |
format_sse_message(data)
async
Format the data as a proper SSE message
Source code in flowfile_core/flowfile_core/routes/logs.py
30 31 32 | |
stream_logs(flow_id, idle_timeout=300, current_user=Depends(get_current_user_from_query))
async
Streams logs for a given flow_id using Server-Sent Events. Requires authentication via token in query parameter. The connection will close gracefully if the server shuts down.
Source code in flowfile_core/flowfile_core/routes/logs.py
109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 | |
public
flowfile_core.routes.public
Classes:
| Name | Description |
|---|---|
GeneratedKey |
Response model for the generate key endpoint. |
SetupStatus |
Response model for the setup status endpoint. |
Functions:
| Name | Description |
|---|---|
docs_redirect |
Redirects to the documentation page. |
generate_key |
Generate a new master encryption key. |
get_setup_status |
Get the current setup status of the application. |
GeneratedKey
pydantic-model
Bases: BaseModel
Response model for the generate key endpoint.
Show JSON schema:
{
"description": "Response model for the generate key endpoint.",
"properties": {
"key": {
"title": "Key",
"type": "string"
},
"instructions": {
"title": "Instructions",
"type": "string"
}
},
"required": [
"key",
"instructions"
],
"title": "GeneratedKey",
"type": "object"
}
Fields:
-
key(str) -
instructions(str)
Source code in flowfile_core/flowfile_core/routes/public.py
25 26 27 28 29 | |
SetupStatus
pydantic-model
Bases: BaseModel
Response model for the setup status endpoint.
Show JSON schema:
{
"description": "Response model for the setup status endpoint.",
"properties": {
"setup_required": {
"title": "Setup Required",
"type": "boolean"
},
"master_key_configured": {
"title": "Master Key Configured",
"type": "boolean"
},
"mode": {
"title": "Mode",
"type": "string"
},
"projects_enabled": {
"title": "Projects Enabled",
"type": "boolean"
},
"projects_confined": {
"title": "Projects Confined",
"type": "boolean"
},
"git_available": {
"title": "Git Available",
"type": "boolean"
}
},
"required": [
"setup_required",
"master_key_configured",
"mode",
"projects_enabled",
"projects_confined",
"git_available"
],
"title": "SetupStatus",
"type": "object"
}
Fields:
-
setup_required(bool) -
master_key_configured(bool) -
mode(str) -
projects_enabled(bool) -
projects_confined(bool) -
git_available(bool)
Source code in flowfile_core/flowfile_core/routes/public.py
14 15 16 17 18 19 20 21 22 | |
docs_redirect()
async
Redirects to the documentation page.
Source code in flowfile_core/flowfile_core/routes/public.py
32 33 34 35 | |
generate_key()
async
Generate a new master encryption key.
Source code in flowfile_core/flowfile_core/routes/public.py
57 58 59 60 61 62 63 64 65 | |
get_setup_status()
async
Get the current setup status of the application.
Source code in flowfile_core/flowfile_core/routes/public.py
38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 | |
secrets
flowfile_core.routes.secrets
Manages CRUD (Create, Read, Update, Delete) operations for secrets.
This router provides secure endpoints for creating, retrieving, and deleting sensitive credentials for the authenticated user. Secrets are encrypted before being stored and are associated with the user's ID.
Functions:
| Name | Description |
|---|---|
create_secret |
Creates a new secret for the authenticated user. |
delete_secret |
Deletes a secret by name for the authenticated user. |
get_secret |
Retrieves a specific secret by name for the authenticated user. |
get_secrets |
Retrieves all secret names for the currently authenticated user. |
create_secret(secret, current_user=Depends(get_current_active_user), db=Depends(get_db))
async
Creates a new secret for the authenticated user.
The secret value is encrypted before being stored in the database. A secret name must be unique for a given user.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
secret
|
SecretInput
|
A |
required |
current_user
|
The authenticated user object, injected by FastAPI. |
Depends(get_current_active_user)
|
|
db
|
Session
|
The database session, injected by FastAPI. |
Depends(get_db)
|
Raises:
| Type | Description |
|---|---|
HTTPException
|
400 if a secret with the same name already exists for the user. |
Returns:
| Type | Description |
|---|---|
Secret
|
A |
Source code in flowfile_core/flowfile_core/routes/secrets.py
77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 | |
delete_secret(secret_name, current_user=Depends(get_current_active_user), db=Depends(get_db))
async
Deletes a secret by name for the authenticated user.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
secret_name
|
str
|
The name of the secret to delete. |
required |
current_user
|
The authenticated user object, injected by FastAPI. |
Depends(get_current_active_user)
|
|
db
|
Session
|
The database session, injected by FastAPI. |
Depends(get_db)
|
Returns:
| Type | Description |
|---|---|
None
|
An empty response with a 204 No Content status code upon success. |
Source code in flowfile_core/flowfile_core/routes/secrets.py
166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 | |
get_secret(secret_name, current_user=Depends(get_current_active_user), db=Depends(get_db))
async
Retrieves a specific secret by name for the authenticated user.
Note: This endpoint returns the secret name and metadata but does not expose the decrypted secret value.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
secret_name
|
str
|
The name of the secret to retrieve. |
required |
current_user
|
The authenticated user object, injected by FastAPI. |
Depends(get_current_active_user)
|
|
db
|
Session
|
The database session, injected by FastAPI. |
Depends(get_db)
|
Raises:
| Type | Description |
|---|---|
HTTPException
|
404 if the secret is not found. |
Returns:
| Type | Description |
|---|---|
Secret
|
A |
Source code in flowfile_core/flowfile_core/routes/secrets.py
119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 | |
get_secrets(current_user=Depends(get_current_active_user), db=Depends(get_db))
async
Retrieves all secret names for the currently authenticated user.
Note: This endpoint returns the secret names and metadata but does not expose the decrypted secret values.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
current_user
|
The authenticated user object, injected by FastAPI. |
Depends(get_current_active_user)
|
|
db
|
Session
|
The database session, injected by FastAPI. |
Depends(get_db)
|
Returns:
| Type | Description |
|---|---|
|
A list of |
Source code in flowfile_core/flowfile_core/routes/secrets.py
43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 | |