macro_gen Dev Log #1: Teaching a Computer to Size an Inverter, Honestly

Part 1 of an ongoing series on macro_gen, an automated cell-characterization tool. Why automate this space, how today’s sizing loop actually works in Rust, real ngspice results from varying Vinv, and where evolutionary search, game theory, and GNNs come in next.

This is the first entry in a developer log for macro_gen, a Rust tool that drives ngspice in-process to characterize and size standard cells against a PDK. I’m writing this series as the tool grows instead of after the fact, which means part 1 is deliberately small: one cell (an inverter), one objective (switching threshold), one PDK (SKY130). Everything below is real, I ran it this afternoon against the actual SKY130 models, not a mocked-up example. The honesty is the point. If I oversell where this is today, I won’t notice where it needs to go.

Why automate this at all

Sizing a single inverter by hand, picking a Wp/Wn ratio, running a DC sweep in ngspice, checking the switching threshold, adjusting, re-running, is a ten-minute task for someone who knows what they’re doing. The problem is that “a cell” is never the unit of work in real design. You need this inverter at multiple switching thresholds because downstream logic has different noise-margin requirements. You need it across process corners. You need it at multiple channel lengths for area/leakage trade-offs. You need NAND, NOR, and eventually sequential cells at all of the above. The moment you multiply cell types by thresholds by corners by lengths by PDKs, “ten minutes by hand” becomes weeks of repetitive simulation, and weeks of repetitive simulation is exactly the kind of thing a human should not be doing manually, both because it’s slow and because it’s the kind of task where fatigue introduces the transcription errors that hand-characterized libraries are notorious for.

But the deeper motivation isn’t just “automate the tedious part.” It’s that once sizing is a function you can call, characterize(cell, target, corner, geometry) -> (sizing, ppa_estimate), the output of characterization stops being a one-off spice deck and becomes a dataset. A dataset you can query: how does drive strength trade off against area as I move the switching threshold? How sensitive is a cell’s delay to a 10% change in channel length at this corner? Those are exactly the questions an architecture-level design-space exploration needs answered, not once, but thousands of times, while it searches for where to spend area and power budget across a chip. macro_gen’s real job, eventually, isn’t “generate a deck.” It’s “turn characterization into a queryable surface that architectural optimization can search over.” Today it does the first half of that sentence.

How today’s sizing loop actually works

The pipeline is a straight line: load config, validate it, extract device parameters from ngspice, compute an analytical seed, sweep around that seed in simulation, keep the closest match. Here’s the config schema that describes one run:

#[derive(Debug, Deserialize, Clone)]
#[serde(deny_unknown_fields)]
pub struct ReferenceInverter {
    pub name: String,
    pub nmos_w: f64,
    pub nmos_l: f64,
    #[serde(default)]
    pub pmos_l: Option<f64>,
    pub inverter_threshold: f64,
    #[serde(default)]
    pub wl_sweep_granularity: Option<f64>,
    #[serde(default)]
    pub wl_sweep_sample_count: Option<u32>,
    #[serde(default = "default_false")]
    pub w_is_multiple_of_w_min: bool,
    #[serde(default = "default_false")]
    pub l_is_multiple_of_l_min: bool,
}

inverter_threshold is the target Vinv, the input voltage at which Vout crosses Vout=Vin on a DC sweep. That’s the one number today’s tool is trying to hit.

The analytical piece is small and deliberately modest. Rather than deriving Wp/Wn from a textbook long-channel square-law formula, it measures the NMOS and PMOS drain currents at a shared characterization bias and takes their ratio:

/// Returns Wp/Wn as the ratio of the two devices' simulated drain currents
/// (`id`) at the normalize stage's characterization bias — both devices
/// share the same W/L there, so the ratio is meaningful without needing
/// to know it.
///
/// Deliberately uses measured `id`, not a reconstruction from `u0`/`cox`:
/// raw `u0` is a per-bin BSIM fit constant, not the true effective
/// mobility, and produces badly wrong ratios. This is only a seed for
/// `sweep_wp_for_target_vth`, which measures the real switching threshold
/// directly, so it only needs to be in the right ballpark.
pub fn id_based_w_l_n_p_ratio(symbols: &SymbolTable) -> f64 {
    let id_n = symbols.get("nmos.id").unwrap_or(0.0).abs();
    let id_p = symbols.get("pmos.id").unwrap_or(0.0).abs();
    id_n / id_p
}

Then the actual refinement is brute force, on purpose: regenerate and re-simulate the whole deck for a fixed grid of candidate Wp values centered on that seed, and keep whichever one’s simulated switching threshold lands closest to the target.

fn sweep_wp_for_target_vth(
    config: &Config,
    project_dir: &Path,
    session: &NgspiceSession,
    wn: f64, ln: f64, lp: f64, wp_seed: f64,
) -> Result<(f64, f64), GenerateError> {
    let ri = &config.reference_inverter;
    let step = ri.wl_sweep_granularity.unwrap_or(config.environment.min_width);
    let sample_count = ri.wl_sweep_sample_count.unwrap_or(DEFAULT_SWEEP_SAMPLE_COUNT);
    let half = (sample_count / 2) as i32;
    let target = ri.inverter_threshold;

    let mut best: Option<(f64, f64, f64)> = None; // (wp, vth, |diff|)
    for i in -half..=half {
        let candidate = (wp_seed + f64::from(i) * step).max(config.environment.min_width);
        let measured = measure_switching_threshold(config, project_dir, session, wn, ln, candidate, lp)?;
        let diff = (measured - target).abs();
        if best.is_none_or(|(_, _, d)| diff < d) {
            best = Some((candidate, measured, diff));
        }
    }
    best.map(|(wp, vth, _)| (wp, vth)).ok_or(/* ... */)
}

Each candidate is a full ngspice run: render a fresh deck, source it, run a DC sweep, measure the crossing. Nothing is cached across candidates. That’s expensive, and it’s expensive for a specific, honest reason logged right in the code:

Regenerates and re-sources the whole deck per candidate rather than altering a live device’s W: this PDK’s subckt-wrapped MOSFETs don’t support in-place width alteration.

That’s SKY130 being SKY130, and it’s the kind of workaround that only shows up once you actually try the “fast” path and watch it fail.

The analytical model, and how far it actually is from simulation

Here’s the part I want to be most honest about, because it’s easy to undersell. The “analytical model” in macro_gen today does not predict switching threshold from sizing. It has no formula for Vinv(Wp, Wn, device params) at all. All it does is estimate a reasonable starting ratio from measured currents, threshold-independent, the same seed regardless of what Vinv you’re targeting. Every actual Vinv number in the tool’s output comes from simulation, not the model. The “analytical” layer earns its keep purely by picking a good starting point for a search that is otherwise entirely empirical.

I wanted to know exactly how good that starting point is, so I ran the sweep at five different target thresholds, holding geometry fixed (Wn = Ln = Lp = 0.15µm minimum length, Wn = 0.42µm minimum width, ratio-based seed, 11-sample sweep), against the real SKY130 tt corner models:

Target Vinv (V) Seed Wp (µm) Final Wp (µm) Measured Vinv (V) Error (V)
0.70 1.1269 0.4200 0.8307 +0.1307
0.80 1.1269 0.4200 0.8307 +0.0307
0.90 1.1269 1.1269 0.8994 −0.0006
1.00 1.1269 3.2269 0.9507 −0.0493
1.10 1.1269 3.2269 0.9507 −0.1493

The seed Wp is identical in every row, 1.1269µm, because the analytical layer never looks at the target. That’s the honest gap: near 0.9V the seed happens to sit almost exactly on the answer (−0.6mV error, better than the sweep’s own granularity would suggest is possible), because 0.9V is close to what a current-ratio-matched inverter naturally produces on this process. Move away from that point and the model has nothing to say, the tool falls back entirely on a fixed-width grid search that, at both ends of this table, ran out of room: the sweep window is ±5 steps of 0.42µm around the seed, so it floors at Wn’s width (0.42µm) for low targets and ceilings at 3.2269µm for high ones, and simply can’t reach targets outside that window no matter how far off they are. Each of these five runs took ~36 seconds and 11 full ngspice simulations to tell me that.

That fixed-width, seed-blind window is a real limitation, not a rounding error, and it’s exactly the kind of thing that’s invisible until you push the tool outside the one design point (0.9V, SKY130, minimum length) it was implicitly built around.

Every number in that table is real, not simplified for the post, so don’t take my word for it:

Reproduction bundle
macro-gen-dev-log-01-configs.tar.gz
5 TOML configs (one per target Vinv) + a README with exact repro steps · 1.8 KB

What this buys us for architecture-level optimization

Even this small, it’s already useful for more than one Vinv. Each row in that table is a (target, sizing, achieved-result, cost-in-simulator-time) tuple, and a table like that, generated at scale across thresholds, corners, and geometries, is the raw material for the actual question I care about: given an area or power budget, which cell configurations sit on the Pareto frontier, and how does that frontier shift as you change process assumptions? That’s an architectural question, not a circuit one, but you can’t answer it without circuit-accurate characterization data feeding it. The five rows above are a toy version of exactly that data, and today it costs three minutes of wall clock for five points on one axis with everything else held fixed. That cost is the whole reason the next section exists.

Where this has to go: smarter search, not just more brute force

Eleven full SPICE runs per design point is fine for a demo and would not survive contact with a real sweep: five thresholds × three lengths × three corners × two PDKs is 90 points, roughly 54 minutes of wall clock for one cell type, and that number only grows once NAND, NOR, and sequential cells join the grid. Brute-force grid search doesn’t scale, and it isn’t even a good use of the simulations it does run, most candidates in the table above told me nothing I couldn’t have guessed from their neighbors.

A few directions I want to try, in roughly the order I’d reach for them:

Better search before better models. Bayesian optimization or a small evolutionary algorithm over the Wp/geometry space would spend simulation budget where the objective function is actually uncertain, instead of a fixed grid that wastes runs on candidates far from the target (like the 0.42µm floor above, hit four times in a row for no new information after the first). This alone should cut the simulation count per design point by a large factor without touching accuracy.

Multi-objective search once there’s more than one objective. Right now there’s exactly one target, switching threshold. The moment delay, leakage, and area join it, picking a single “best” candidate stops making sense, you want a Pareto set, not a point. That’s a natural fit for evolutionary multi-objective algorithms (NSGA-II and similar), and it’s also where game-theoretic framing gets interesting: sizing decisions for cells that share a rail or a critical path aren’t independent, they’re closer to a coordination game between cells competing for the same power/delay budget, and treating library-wide sizing as a single joint optimization rather than cell-by-cell independent sweeps is the more honest formulation of the actual problem.

Learned surrogates once there’s enough data to learn from. A DNN or a simple ANN trained on (geometry, device params, corner) → (Vinv, delay, leakage) could replace most of the sweep with a cheap forward pass, falling back to real simulation only near the predicted optimum to verify it. And because gate topologies are naturally graphs, transistors as nodes, connections as edges, a GNN is the more principled version of that idea: a surrogate that generalizes across cell topologies (inverter → NAND → NOR → more complex gates) instead of needing to be retrained from scratch for each one. That’s further out than the search-algorithm work, it needs a real dataset to train on first, which is exactly what a better search strategy would start producing faster.

None of this is built yet. This entry is the “here’s where we are and here’s the honest gap” post. The next one will be either the evolutionary search replacing the fixed grid, or a second cell type proving out whether today’s architecture (a single hardcoded inverter module) can actually support one without a rewrite. Probably in that order.