Code generation¶
To evaluate a Function on concrete numerical values, scaly generates C code
from the symbolic computational graph. This happens both in Python and at the
user's request for deployment in a C or C++ application. This allows, for
example, a controller in a ROS node to reuse the optimal controller written
during prototyping without having to rewrite it. The model is written once,
in Python, and both uses compile from it.
This page assumes familiarity with scaly functions and basic C concepts (such as pointers and how to compile C code).
Exported source and compiled functions¶
Consider a function representing \(E(x)=\sum_i x_i^2\):
import scaly as sc
@sc.function(sc.L("x", 3), sc.L("energy", ...))
def energy(x: sc.Expr) -> sc.Expr:
return sc.sumsqr(x)
The decorated energy is an sc.Function. Constructing it records the
expression graph without generating or compiling C. There are two ways to
use that graph:
- A numerical Python call automatically generates C, compiles a shared library, and evaluates it. Scaly caches the library for later calls. This is the just-in-time, or JIT, compilation path.
- Export the
Functionto source and header files in a specific directory and integrate them with your application's build system. This is the ahead-of-time, or AOT, compilation path.
Both paths use the same code generator. Exporting source does not compile it, and exporting alone does not directly require numerical inputs or a C compiler.
from pathlib import Path
from scaly.codegen import write_module
module = write_module(energy, Path("generated"))
print(module.header_name, module.source_name)
# energy.h energy.c
write_module writes generated/energy.h and generated/energy.c and returns
a CModule. This object holds the generated text, filenames, workspace size,
the build recipe, and link flags. Existing files with those names are
overwritten.
The command-line equivalent refers to the Python object by module and attribute
name. If energy is defined in an importable module named model, the command is:
model:energy means import model and use its energy attribute. The module
must be importable in the command's Python environment. A model.py in the
current directory is one way to provide it.
The C header¶
The generated header declares one struct per input and output, a workspace
struct, and an energy_call helper that evaluates the function. These are
excerpts from energy.h:
typedef struct { SCALY_ALIGNAS(16) double data[3]; } energy_x_t;
typedef struct { SCALY_ALIGNAS(16) double data[1]; } energy_energy__t;
typedef struct { SCALY_ALIGNAS(16) double data[energy_SZ_W > 0 ? energy_SZ_W : 1]; } energy_workspace_t;
static inline int energy_call(const energy_x_t* x, energy_energy__t* energy_, energy_workspace_t* workspace);
Each struct holds one contiguous buffer of double values. Arrays use
row-major order, as in a contiguous NumPy array, and a scalar occupies one
entry. Python tuple groups do not become extra buffers. The output name becomes
energy_ inside the generated interface to avoid a collision with the function
name, which accounts for the double underscore in energy_energy__t.
The structs own their storage. You can allocate them on the stack, statically,
or on the heap. C11 and C++ buffers request 16-byte alignment, while C99 uses
natural alignment. The workspace struct reserves at least one entry because
standard C does not allow a zero-length array. The header reports the real
requirement in doubles as energy_SZ_W.
For the energy model above, the following complete C caller supplies one
three-element input and reads one scalar output. The compilation command below
uses main.c as its filename:
#include <stdio.h>
#include "energy.h"
int main(void) {
energy_x_t x = {{1.0, 2.0, 3.0}};
energy_energy__t result;
energy_workspace_t workspace;
int status = energy_call(&x, &result, &workspace);
if (status != 0) return status;
printf("%g\n", result.data[0]);
return 0;
}
The program prints 14. It links the C math library with -lm and needs no
Python runtime. The interface reference
lists the argument checks and return codes.
The pointer entry¶
energy_call is a header-only helper. The compiled symbol underneath it is a
pointer-based entry point with the same shape for every generated function:
arg contains one pointer per input and res one pointer per output, in
declaration order. w is the workspace of energy_SZ_W doubles, or NULL
when that size is zero. iw and mem are unused and take NULL and 0.
The CasADi section explains why they exist.
The helper only fills the two pointer arrays and forwards to this entry.
Call the entry directly in two situations. The first is memory that another
library already owns, such as a message array in a ROS node or an Eigen
matrix, which the typed structs would require copying. The second is loading
the compiled library with dlopen and no header, where this fixed signature
is the whole contract. The C++ header and the CasADi layer are both built on
the same symbol, so every generated module can be called this way.
Generated text and export options¶
render_c_module returns a CModule without writing files:
from scaly.codegen import render_c_module
module = render_c_module(energy)
print(module.workspace_size) # 0
print(module.link_flags) # ()
header = module.header # C declarations and buffer metadata
source = module.source # C implementation
The zero workspace size means this function needs no caller-provided scratch
buffer. Its local C variables still occupy memory. Empty link_flags means
there are no additional solver or vector-math libraries to link beyond -lm.
The main export choices are:
| Python keyword | Command-line option | Result |
|---|---|---|
lang="c" |
--lang c |
C header with typed buffer structs, the default |
lang="cpp" |
--lang cpp |
C++ header with buffer types and a namespace |
casadi=True |
--casadi |
Metadata functions for compatible external-function loaders |
C++ buffers¶
lang="cpp" generates a C++17 header. The implementation remains C, which
lets the C and C++ interfaces use the same compiled numerical function:
cpp_module = write_module(energy, Path("generated_cpp"), lang="cpp")
print(cpp_module.header_name, cpp_module.source_name)
# energy.hpp energy.c
The header contains a Buffer template with inline storage and compile-time
shape information. Its declaration begins as follows, with indexing methods
omitted from this excerpt:
template <typename T, std::size_t... Ns>
struct alignas(16) Buffer {
static constexpr std::size_t ndim = sizeof...(Ns);
static constexpr std::array<std::size_t, ndim> shape = {Ns...};
static constexpr std::size_t size = (Ns * ... * std::size_t{1});
T data[size > 0 ? size : 1];
T* ptr() { return data; }
const T* ptr() const { return data; }
// Indexing methods omitted.
};
Buffer<double, 2, 3> holds a two-by-three matrix in row-major order.
buffer(i, j) accesses an element, with bounds assertions in debug builds.
Buffer<double> is a scalar, accessed as buffer(). Empty buffers reserve
one physical element while reporting a logical size of zero.
Each function has its own namespace. For energy, the generated aliases are:
using x_t = Buffer<double, 3>;
using energy__t = Buffer<double>;
using workspace_t = Buffer<double, energy_SZ_W>;
These aliases live inside namespace energy. The output alias contains two
underscores for the same name-collision reason as the typed C struct.
A complete main.cpp using this header is:
#include <iostream>
#include "energy.hpp"
int main() {
energy::x_t x = {{1.0, 2.0, 3.0}};
energy::energy__t result{};
energy::workspace_t workspace{};
int status = energy::call(x, result, workspace);
if (status != 0) return status;
std::cout << result() << '\n';
}
cc -O2 -c generated_cpp/energy.c -o energy.o
c++ -std=c++17 -Igenerated_cpp main.cpp energy.o -lm -o energy_cpp_demo
./energy_cpp_demo
This prints 14. The energy::call helper constructs pointer arrays and calls
the C entry point, declared with extern "C" linkage. Buffer owns its data.
To use arrays already owned by another library without copying them into a
Buffer, call the pointer interface directly.
Sparse output patterns¶
Sparse derivative functions write compact value arrays. The generated header also contains fixed coordinate tables so the caller can identify the matrix entry corresponding to each value. For example:
@sc.function(sc.L("x", 3), sc.L("y", ...))
def measurements(x: sc.Expr) -> sc.Expr:
return sc.stack([x[0] * x[2], x[1], x[0] + x[2]])
measurements_jac = sc.sparse_jacobian(measurements, "y", "x", name="measurements_jac")
sparse_module = write_module(measurements_jac, Path("generated_sparse"))
The Jacobian is
Its output has five values instead of nine. The following declarations are
emitted in measurements_jac.h:
#define measurements_jac_spjac_y_x_NNZ 5
#define measurements_jac_spjac_y_x_NROW 3
#define measurements_jac_spjac_y_x_NCOL 3
static const int measurements_jac_spjac_y_x_rows[5] = {0, 0, 1, 2, 2};
static const int measurements_jac_spjac_y_x_cols[5] = {0, 2, 1, 0, 2};
static const int measurements_jac_spjac_y_x_csr_row_ptr[4] = {0, 2, 3, 5};
static const int measurements_jac_spjac_y_x_csr_col_ind[5] = {0, 2, 1, 0, 2};
static const int measurements_jac_spjac_y_x_csr_val_perm[5] = {0, 1, 2, 3, 4};
static const int measurements_jac_spjac_y_x_csc_col_ptr[4] = {0, 2, 3, 5};
static const int measurements_jac_spjac_y_x_csc_row_ind[5] = {0, 2, 1, 0, 2};
static const int measurements_jac_spjac_y_x_csc_val_perm[5] = {0, 3, 2, 1, 4};
The prefix combines the function name and derivative output name. For input
[2, 3, 4], the value array is [4, 2, 1, 1, 1]. Entry k belongs at
(rows[k], cols[k]). Coordinates are not guaranteed to be sorted, even though
the row order happens to be sorted in this example.
Compressed sparse row and compressed sparse column formats, or CSR and CSC,
need their own value orders. Their _val_perm arrays give the conversion:
values_csc[k] = values[csc_val_perm[k]]. Here the CSC values become
[4, 1, 1, 2, 1]. The CSR values already have the required order.
C++ headers expose the same metadata as constexpr arrays inside a namespace
named after the derivative output. The sparsity guide explains
the corresponding Python patterns and triangle choices for Hessians.
CasADi-compatible exports¶
The pointer entry takes the same five arguments, in the
same order, as a function generated by CasADi 3.8, including its
integer-workspace argument iw and its memory-handle argument mem. Scaly
functions are stateless, so both are ignored. The one difference in the
signature is iw, which Scaly declares as int* where CasADi uses
casadi_int*. Because the entry never reads it, CasADi and acados can pass
their own pointer unchanged. The signature alone does not tell a loader how many inputs a
function has or how large its buffers are. casadi=True adds the metadata
queries of CasADi's external-function interface:
For this model, the additional header declarations include:
casadi_int energy_n_in(void);
casadi_int energy_n_out(void);
const char* energy_name_in(casadi_int i);
const char* energy_name_out(casadi_int i);
const casadi_int* energy_sparsity_in(casadi_int i);
const casadi_int* energy_sparsity_out(casadi_int i);
int energy_work(casadi_int* sz_arg, casadi_int* sz_res, casadi_int* sz_iw, casadi_int* sz_w);
int energy_checkout(void);
void energy_release(int mem);
energy_n_in() and energy_n_out() both return one. The name queries return
"x" and "energy" at index zero. The sparsity queries describe a dense
three-by-one input and one-by-one output. energy_work returns the required
pointer-array and workspace sizes. energy_checkout() returns zero and
energy_release does nothing because scaly does not use memory handles.
For example, on Linux this command builds a shared library for CasADi:
With CasADi installed, it can then be loaded and called:
import casadi as ca
external_energy = ca.external("energy", "generated_casadi/libenergy.so")
print(float(external_energy([1.0, 2.0, 3.0]))) # 14.0
The generated casadi_int macro defaults to long long int, matching CasADi's
loader. Consumers expecting another integer width need a matching build.
In particular, acados uses -Dcasadi_int=int. The option works with either
lang="c" or lang="cpp".
For sparse outputs, this mode writes values in CSC order. If native scaly ordering differs, the generated entry point permutes the values using extra workspace. The module's workspace size includes that storage, and its header tables describe the converted output order.
Buffer compatibility
Dense matrices with both dimensions greater than one, and dense arrays of
rank greater than two, are rejected by casadi=True. Scaly uses row-major
storage while CasADi uses column-major storage. Scalars, vectors, and compact
sparse outputs are supported. Every input and output pointer is still
required, unlike CasADi's convention of treating null inputs as zero and
skipping null outputs.
Solver export and deployment dependencies¶
write_module and scaly_codegen accept an sc.Solver as well as a
Function. They export the solver's underlying solve.function, including
its objective, constraints, derivatives, and native solver wrapper.
render_c_module and workspace_size take a Function, so those calls use
solve.function explicitly.
The C interface always requires the initial variables, initial multipliers,
and parameters. Defaults supplied by the Python Solver wrapper do not
become defaults in C. Generated solver wrappers also keep static state, so
the same wrapper cannot be called concurrently.
Ordinary generated model code has no dependency on Python, NumPy, or SciPy.
Solver code additionally needs the corresponding native solver headers and
libraries. The CModule.link_flags property reports the flags discovered for
those dependencies and the selected math library in the current environment. Reading it requires the
libraries to be available, even though rendering their wrapper source does not.
For deployment on another machine, provide libraries and paths appropriate to
that target. See Solver backends.
CPU targets, vector lanes, and math libraries¶
A generated module carries a BuildRecipe in module.recipe. It records the
CPU target, lane width, C dialect, and numerical options used to generate the
source. The header and source begin with a comment showing matching GCC and
Clang compilation commands and required math libraries.
A lane is one element processed alongside other elements in a vector operation. Scaly can group independent loop iterations into lanes without changing the function's input and output layout. For example:
vector_module = render_c_module(
energy, cpu="x86-64-v3", lanes=4, dialect="gnu", vector_libm="glibc",
)
print(vector_module.recipe.cpu_flags) # ('-march=x86-64-v3',)
print(vector_module.link_flags) # ('-lmvec',)
These options request a target configuration. They do not guarantee that a particular small model uses vector instructions. The compiler widens eligible loops and can reduce their width to limit register pressure.
CPU baseline and lane width¶
Export defaults to cpu="generic", which adds no target-specific compiler
flags, and lanes="auto". Automatic widths follow the target macros set by
the C compiler:
| Target features | Requested lanes |
|---|---|
| AVX-512 | 8 |
| AVX or fixed Arm SVE width of at least 256 bits | 4 |
| SSE2 or AArch64 | 2 |
| Other targets | 1 |
A fixed lanes=1, 2, 4, or 8 bypasses this detection. One lane disables
widening. With automatic selection, -DSCALY_LANES=4 can select the width at
C compilation time. Partial groups use guarded stores, and input and output
arrays need no extra padding. Workspace capacity supports up to eight lanes,
so changing the lane selection does not change the header's workspace size.
cpu="native" selects -march=native on x86 or -mcpu=native on AArch64.
The other named targets are x86-64-v3, x86-64-v4, and apple-m4.
module.recipe.cpu_flags gives their matching flags.
The target belongs to the application build
Export records the requested target without checking the deployment machine.
Compile with the recipe's flags and choose a baseline supported by every
destination. A native binary is intended for the build machine, rather
than an arbitrary computer with the same operating system.
C dialect and vector math¶
dialect="gnu" is the default and uses GNU vector extensions supported by
GCC and Clang. dialect="c" emits scalar lane loops in C99 syntax for
compilers without those extensions. This option is separate from lang,
which chooses the language of the caller's header.
Export defaults to vector_libm="none", so each lane uses scalar math calls.
vector_libm="glibc" selects supported libmvec implementations of sin,
cos, tan, exp, log, pow, and tanh. It requires the GNU dialect,
x86-64 with compatible target features, and glibc. tanh additionally
requires glibc 2.35 or newer. Generated guards reject incompatible builds.
The recipe adds -lmvec, in addition to the usual -lm math-library link.
Vector math can produce different rounded results from scalar math. Likewise,
reciprocal=True allows invariant division x / y to become multiplication
by a precomputed 1 / y, which changes rounding and can change overflow.
Both are separate numerical choices. Neither enables general fast-math flags.
See arithmetic semantics
for simplifications that apply independently of these options.
The same choices are available to scaly_codegen as --cpu, --lanes,
--dialect, --vector-libm, and --reciprocal. The
render-options reference lists their values.
Compilation and caching in Python¶
A numerical call compiles on first use unless a matching artifact already
exists in the disk cache. Scaly then keeps the loaded library in memory.
The Python compiler detects a native CPU recipe with a fixed lane width.
On supported glibc x86-64 hosts it also selects vector math automatically.
SCALY_VECTOR_LIBM=none forces scalar math calls, as described in
environment settings.
The cache key includes the generated source and its recipe, function name,
interface version, and compiler and link flags. New numerical inputs do not
cause recompilation. The key does not identify the compiler executable or the
exact CPU, so environment settings
describes when to clear the cache.
compile() prepares the library without evaluating the function:
This is useful before a time-sensitive calculation. It also reuses a cached
library when available. For a solver, solve.function.compile() prepares the
compiled wrapper without running an optimization.
The default cache directory is $XDG_CACHE_HOME/scaly/jit, or
~/.cache/scaly/jit when XDG_CACHE_HOME is unset. SCALY_CACHE_DIR overrides
it. Deleting cached files does not invalidate an already loaded function, but
future compilations rebuild missing artifacts.
Despite its name, recompile() clears the function's compiled handle and disk
cache entry without immediately compiling. The next numerical call or explicit
compile() prepares the library again.
Scaly does not fall back to Python evaluation when compilation fails.
uv run scaly_toolchain reports the compiler, cache directory, and resolved
solver libraries when diagnosing such a failure.
Working memory¶
workspace_size obtains the caller-provided workspace requirement without
rendering all the generated text:
The number counts double entries in w, not bytes or total temporary memory.
For an export with non-default options, use that module's workspace_size or
its header's <name>_SZ_W constant. In particular, CasADi sparse-output
conversion can require additional workspace.