macro_gen Dev Log #2: Sizing Combinational Logic by Logical Effort
macro_gen now sizes arbitrary combinational CMOS logic by logical effort, relative to the inverter characterized in part 1, and buffers outputs to log4(F) stages. Method with a fully worked example, calibration of the model constants, and a comparison against SPICE on eleven SKY130 benchmarks.
Summary. We taught a computer to size transistors
for combinational logic. Given a design in the hw and
comb dialects of CIRCT [1], [2] and a PDK, macro_gen (v0.2.0)
builds every gate as a static CMOS pull-up and pull-down network, sizes
each transistor by the method of logical effort [3] relative to the reference inverter
characterized in part 1, and
inserts inverters so each output path has the delay-optimal number of
stages. We calibrated the model’s constants in ngspice [4] and compared its delay estimates
against SPICE on eleven benchmarks in SKY130 [5]. For every buffered netlist the
estimate is 5% to 24% below the measured delay. For unbuffered netlists
the estimate is within 24% on eight of eleven designs and overestimates
by 45% to 147% on the other three. In SPICE, buffering reduces the worst
case delay by 1.3× to 3.8× on eight designs, leaves two unchanged, and
makes one 11% slower. The worst case, an 8 input AND built as a linear
chain, comes within 13% when the same function is built as a balanced
tree, which is also 2.3× faster.
1. Problem
Part 1 sized one inverter to a target switching threshold by simulation: an 11 point ngspice sweep, about 50 s per design point. Repeating that for every gate type does not scale. Logical effort avoids it: once one inverter is characterized, every other gate can be sized analytically relative to it. This entry applies that method automatically to arbitrary combinational netlists and measures how well it predicts delay.
2. Model
The delay of a logic stage, in units of a process constant \(\tau\), is
\[ d = f + p, \qquad f = g\,h, \qquad h = \frac{C_\text{out}}{C_\text{in}}, \]
where \(g\) is the stage’s logical effort (its input capacitance relative to an inverter delivering the same current), \(h\) its electrical effort, and \(p\) its parasitic delay. Along a path of \(N\) stages, the path effort is
\[ F = G\,B\,H, \qquad G = \prod_i g_i, \qquad B = \prod_i b_i, \qquad H = \frac{C_\text{load}}{C_{\text{in},1}}, \]
where \(b_i = (C_\text{on path} + C_\text{off path}) / C_\text{on path}\) is the branching effort at stage \(i\). Delay is minimized when every stage carries the same effort, and the optimal number of stages is
\[ \hat f = F^{1/N}, \qquad \hat N = \operatorname{round}\left(\log_\rho F\right), \quad \rho \approx 4. \]
Transistor sizing. The reference inverter has \(W_n = 0.42\,\mu\text{m}\) and \(W_p = 1.1269\,\mu\text{m}\), so \(\gamma = W_p / W_n\). Each gate is sized so its worst pull-up and pull-down paths conduct like the inverter’s: a transistor in a series stack of \(n\) is \(n\) times wider. With widths in units of the inverter’s NMOS, the logical effort of input \(i\) is
\[ g_i = \frac{W_{n,i} + \gamma\,W_{p,i}}{1 + \gamma}. \]
Netlist sizing. Real netlists branch and reconverge, so instead of enumerating paths macro_gen gives every stage the same effort \(f\). In reverse topological order each stage gets drive \(s\) and input capacitance
\[ s = \frac{C_\text{out}}{f}, \qquad C_{\text{in},i} = g_i\, s, \]
where \(C_\text{out}\) is the sum of the input capacitances it drives plus \(C_\text{load}\) at an output port. The largest primary input capacitance falls monotonically as \(f\) rises, so \(f\) is found by bisection such that \(\max C_\text{in} = C_{\text{in},\max}\). On a single path this reproduces \(\hat f = F^{1/N}\).
Buffering. Each output whose path has fewer than \(\hat N\) stages gets \(\hat N - N\) inverters before its port, and the netlist is resized.
3. Using macro_gen
A run is described by one TOML file: the PDK, the reference inverter, the sizing targets and the design. The part that matters here is
[sizing]
cload_cinv = 64.0 # load on each module output port, in C_inv
cin_cinv = 1.0 # largest capacitance any primary input may present
stage_effort = 4.0 # rho, for buffering
[circt]
mlir_path = "c17.mlir"
top_module = "c17"The design is plain CIRCT. c17 [6] has six NAND gates; since
comb has no NAND, each is an AND followed by a NOT, written
as comb.xor with a constant true:
hw.module @c17(in %N1: i1, in %N2: i1, in %N3: i1, in %N6: i1, in %N7: i1,
out N22: i1, out N23: i1) {
%true = hw.constant true
%a10 = comb.and %N1, %N3 : i1
%N10 = comb.xor %a10, %true : i1
...
hw.output %N22, %N23 : i1, i1
}
macro_gen --config benchmarks/c17/c17.toml --emit-verilog --add-bufferThe run log shows every sizing decision, so any number in the netlist
can be traced by hand. For each cell, in signal order: its logical
effort and parasitic delay, the capacitance on each input and on its
output, its drive, and the resulting transistor widths (*
marks a width clamped up to the process minimum):
Sized cells (unbuffered): every stage at f = C_out / s = 14.385, C_in(pin) = g x s
cell | master | g (per input) | p (tau) | C_in (C_inv) | C_out (C_inv) | C_out (fF) | drive s (x inv) | NMOS W (um) | PMOS W (um)
--------+-------------+---------------+---------+---------------+---------------+------------+-----------------+-------------+------------
and2_2 | nand2_x0p10 | 1.272 / 1.272 | 2.000 | 0.133 / 0.133 | 1.500 | 2.23 | 0.104 | 0.42*x2 | 0.42*x2
and2_6 | nand2_x0p39 | 1.272 / 1.272 | 2.000 | 0.500 / 0.500 | 5.657 | 8.43 | 0.393 | 0.42*x2 | 0.44x2
and2_4 | nand2_x0p79 | 1.272 / 1.272 | 2.000 | 1.000 / 1.000 | 11.314 | 16.85 | 0.786 | 0.66x2 | 0.89x2
and2_10 | nand2_x4p45 | 1.272 / 1.272 | 2.000 | 5.657 / 5.657 | 64.000 | 95.34 | 4.449 | 3.74x2 | 5.01x2
and2_0 | nand2_x0p39 | 1.272 / 1.272 | 2.000 | 0.500 / 0.500 | 5.657 | 8.43 | 0.393 | 0.42*x2 | 0.44x2
and2_8 | nand2_x4p45 | 1.272 / 1.272 | 2.000 | 5.657 / 5.657 | 64.000 | 95.34 | 4.449 | 3.74x2 | 5.01x2
For each output, the critical path and every term of the path effort:
Output paths (unbuffered): F = G x B x H, f = F^(1/N), N_opt = log_4 F, delay = N f + P
output | critical path | N | G | B | H | F | f | N_opt | added inv | P (tau) | delay (tau)
-------+--------------------------------------+---+-------+-------+--------+-----------+--------+-------+-----------+---------+------------
N22 | N6 > and2_2 > and2_4 > and2_8 > N22 | 3 | 2.056 | 3.000 | 482.72 | 2976.9555 | 14.385 | 5.77 | 3 | 6.00 | 49.16
N23 | N6 > and2_2 > and2_6 > and2_10 > N23 | 3 | 2.056 | 3.000 | 482.72 | 2976.9555 | 14.385 | 5.77 | 3 | 6.00 | 49.16
The same tables are printed again after buffering, followed by a summary:
Summary: c17
metric | unbuffered | buffered
------------------+------------+---------
stage effort f | 14.386 | 2.982
cells | 6 | 12
transistors | 24 | 36
worst delay (tau) | 49.16 | 26.89
worst delay (FO4) | 9.83 | 5.38
The capacitances in fF use the tool’s own estimate of \(C_\text{inv}\) (1.49 fF, from the gate capacitance at one bias point); Section 5 shows the effective value is 1.93 fF.
Verilog. With --emit-verilog the sized,
buffered design is written as a structural netlist: cell instances and
wires only, no assign and no behavioural code, so it can go
straight to physical design. Each cell master is named for its stage and
drive, nand2_x0p81 is a NAND2 at drive 0.81, and the same
names are used in the SPICE netlist so the two line up for LVS. The
buffering is visible at the end: three inverters of growing size, 2.41,
7.20 and 21.46, in front of each output, and a header comment records
that both outputs are now inverted.
// c17: logical-effort sized structural netlist
// generated by macro_gen 0.2.0 (864ae133bc) on 2026-09-29 17:02:59 UTC
// macro_gen.inverted.N22 = true
// macro_gen.inverted.N23 = true
module c17 (N1, N2, N3, N6, N7, N22, N23);
input N1;
input N2;
input N3;
input N6;
input N7;
output N22;
output N23;
wire N22_buf0_y;
wire N22_buf1_y;
wire N22_pre;
wire N23_buf0_y;
wire N23_buf1_y;
wire N23_pre;
wire net_1;
wire net_3;
wire net_5;
wire net_7;
nand2_x0p35 and2_0 (.in0(N1), .in1(N3), .y(net_1));
nand2_x0p44 and2_2 (.in0(N3), .in1(N6), .y(net_3));
nand2_x0p69 and2_4 (.in0(N2), .in1(net_3), .y(net_5));
nand2_x0p35 and2_6 (.in0(net_3), .in1(N7), .y(net_7));
nand2_x0p81 and2_8 (.in0(net_1), .in1(net_5), .y(N22_pre));
nand2_x0p81 and2_10 (.in0(net_5), .in1(net_7), .y(N23_pre));
inv_x2p41 N22_buf0 (.in0(N22_pre), .y(N22_buf0_y));
inv_x7p20 N22_buf1 (.in0(N22_buf0_y), .y(N22_buf1_y));
inv_x21p46 N22_buf2 (.in0(N22_buf1_y), .y(N22));
inv_x2p41 N23_buf0 (.in0(N23_pre), .y(N23_buf0_y));
inv_x7p20 N23_buf1 (.in0(N23_buf0_y), .y(N23_buf1_y));
inv_x21p46 N23_buf2 (.in0(N23_buf1_y), .y(N23));
endmoduleSPICE. Every master becomes a transistor-level subcircuit. The NAND2 above at drive 0.81 has two series NMOS at \(2 \times 0.81 \times 0.42 = 0.68\,\mu\text{m}\) and two parallel PMOS at \(0.81 \times 1.1269 = 0.91\,\mu\text{m}\):
.subckt nand2_x0p81 in0 in1 y VDD GND
XMN0 y in0 nn0 GND sky130_fd_pr__nfet_01v8 L=0.1500 W=0.6804 nf=1 mult=1
XMN1 nn0 in1 GND GND sky130_fd_pr__nfet_01v8 L=0.1500 W=0.6804 nf=1 mult=1
XMP0 y in0 VDD VDD sky130_fd_pr__pfet_01v8 L=0.1500 W=0.9128 nf=1 mult=1
XMP1 y in1 VDD VDD sky130_fd_pr__pfet_01v8 L=0.1500 W=0.9128 nf=1 mult=1
.ends nand2_x0p81
Report. reports/c17.toml carries the
same results in machine-readable form, with units in every key, for
scripts and design space exploration loops:
[unbuffered]
stage_effort = 14.3855
cells = 6
transistors = 24
delay_tau = 49.1564
delay_fo4 = 9.8313
[buffered]
stage_effort = 2.9821
cells = 12
transistors = 36
delay_tau = 26.8924
delay_fo4 = 5.3785
[[outputs]]
name = "N22"
logic_stages = 3
path_effort = 2976.9556
added_inverters = 3
delay_tau = 49.1564
buffered_delay_tau = 26.89244. Worked example: ISCAS-85 c17
The steps below reproduce the tables in Section 3 by hand. c17 [6] is six NAND2 gates. Take its worst
path, \(N_6 \rightarrow \text{NAND}_\text{a}
\rightarrow \text{NAND}_\text{b} \rightarrow \text{NAND}_\text{c}
\rightarrow N_{22}\) (cells and2_2,
and2_4 and and2_8), with \(C_\text{load} = 64\,C_\text{inv}\) and
\(C_{\text{in},\max} =
1\,C_\text{inv}\).
Step 1: process ratio. \[\gamma = \frac{1.1269}{0.42} = 2.683\]
Step 2: logical effort of a NAND2. Its two NMOS are in series (width 2 each) and its two PMOS in parallel (width 1 each, scaled by \(\gamma\)): \[g_\text{NAND2} = \frac{2 + \gamma}{1 + \gamma} = \frac{4.683}{3.683} = 1.272\]
Step 3: path logical effort. \[G = g_\text{NAND2}^3 = 1.272^3 = 2.056\]
Step 4: branching effort. From the sized netlist, \(\text{NAND}_\text{a}\)’s output drives \(\text{NAND}_\text{b}\) (1.000 \(C_\text{inv}\)) and another NAND off the path (0.500), and \(\text{NAND}_\text{b}\)’s output drives \(\text{NAND}_\text{c}\) and a second NAND of equal size (5.657 each): \[B = \frac{1.000 + 0.500}{1.000} \cdot \frac{5.657 + 5.657}{5.657} = 1.5 \times 2 = 3.0\]
Step 5: electrical effort. The path enters \(\text{NAND}_\text{a}\), whose input presents 0.1326 \(C_\text{inv}\): \[H = \frac{64}{0.1326} = 482.7\]
Step 6: path effort and stage effort. \[F = G\,B\,H = 2.056 \times 3.0 \times 482.7 = 2977, \qquad \hat f = 2977^{1/3} = 14.39\]
Step 7: size the last stage. \(\text{NAND}_\text{c}\) drives the load directly: \[s = \frac{64}{14.39} = 4.449\] \[W_n = 2 \times 4.449 \times 0.42 = 3.74\,\mu\text{m}, \qquad W_p = 1 \times 4.449 \times 1.1269 = 5.01\,\mu\text{m}\]
Step 8: delay. A NAND2’s parasitic delay is \(p = 2\) in this model, so \[d = N(\hat f + p) = 3\,(14.39 + 2) = 49.2\,\tau\]
Step 9: buffering. \[\hat N = \operatorname{round}(\log_4 2977) = \operatorname{round}(5.77) = 6\] Three inverters are added (\(p_\text{inv} = 1\) in the model). After resizing, every stage runs at \(f = 2.98\): \[d = 3\,(2.98 + 2) + 3\,(2.98 + 1) = 26.9\,\tau\]
Step 10: compare with SPICE. With \(\tau = 9.67\,\text{ps}\) from Section 5, the unbuffered estimate is \(49.2 \times 9.67 = 475\,\text{ps}\) against \(410.5\,\text{ps}\) measured (+16%), and the buffered estimate is \(26.9 \times 9.67 = 260\,\text{ps}\) against \(314.9\,\text{ps}\) measured (−17%).
5. Calibration
\(\tau\), \(p_\text{inv}\) and the physical size of \(C_\text{inv}\) were measured, not assumed. The reference inverter drove \(h\) copies of itself, from an input edge shaped by a fanout of four driver, and the average of the rising and falling delays was recorded:
| Fanout | 1 | 2 | 4 | 8 | 16 | 32 | 64 |
|---|---|---|---|---|---|---|---|
| Delay (ps) | 37.5 | 51.0 | 72.9 | 111.9 | 188.4 | 342.2 | 653.3 |
A least squares fit of \(d(h) = \tau\,(p_\text{inv} + h)\) over \(h \ge 4\) (smaller fanouts are distorted by the input edge) gives
\[ \tau = 9.67\,\text{ps}, \qquad p_\text{inv} = 3.51, \qquad d_\text{FO4} = 72.9\,\text{ps}. \]
A capacitor matching the \(h = 64\) delay is 123.4 fF, so the effective input capacitance is \(C_\text{inv} = 1.93\,\text{fF}\) and the 64 \(C_\text{inv}\) load is 123.4 fF.
Two earlier shortcuts were wrong. Taking \(\tau = d_\text{FO4}/5\) assumes \(p_\text{inv} = 1\) and gives 14.6 ps, 50% too high. Estimating \(C_\text{inv}\) from the gate capacitance at a single DC bias point gave 1.49 fF, 23% too low, likely because it misses the Miller contribution of the gate to drain overlap while the inverter switches.
6. Validation against SPICE
Each benchmark was sized with and without buffering, and every
netlist was simulated in ngspice with the SKY130 tt models
at 1.8 V. Each output drove 123.4 fF. Every input that can change an
output was toggled with a 20 ps edge, with the other inputs held at
every combination that lets the change through (found by evaluating the
design’s logic). For each (input, side inputs, output) the delay is the
average of the rising and falling 50% to 50% delays; the worst of these
is the measured delay. The estimate is the model’s delay in \(\tau\) times 9.67 ps.
Structure of the sized netlists:
| Benchmark | Logic stages | \(F\) (worst path) | Inverters added | Transistors |
|---|---|---|---|---|
| inv | 1 | 64 | 2 | 2 → 6 |
| nand2 | 1 | 81.4 | 2 | 4 → 8 |
| aoi21 | 1 | 128 | 3 | 6 → 12 |
| oai22 | 1 | 128 | 3 | 8 → 14 |
| and_or_chain | 2 | 128 | 2 | 8 → 12 |
| maj3 | 2 | 442.5 | 2 | 14 → 18 |
| c17 | 3 | 2977 | 6 | 24 → 36 |
| dec2to4 | 3 | 2316 | 11 | 28 → 50 |
| mux2 | 5 | 655 | 0 | 20 → 20 |
| and8 | 14 | \(2.4 \times 10^{13}\) | 8 | 42 → 58 |
| and8_tree | 6 | 131.6 | 0 | 42 → 42 |
Worst case delay, estimated and measured (ps). Error is (estimate − SPICE) / SPICE; speedup is the measured unbuffered delay over the measured buffered delay.
| Benchmark | Unbuffered est. | SPICE | Error | Buffered est. | SPICE | Error | Speedup |
|---|---|---|---|---|---|---|---|
| inv | 628.4 | 596.5 | +5% | 145.0 | 177.1 | −18% | 3.37× |
| nand2 | 806.1 | 659.6 | +22% | 164.4 | 193.8 | −15% | 3.40× |
| aoi21 | 1273.6 | 879.9 | +45% | 195.1 | 229.1 | −15% | 3.84× |
| oai22 | 1276.2 | 847.2 | +51% | 197.8 | 259.8 | −24% | 3.26× |
| and_or_chain | 264.5 | 299.5 | −12% | 195.1 | 229.1 | −15% | 1.31× |
| maj3 | 488.5 | 469.6 | +4% | 278.5 | 321.0 | −13% | 1.46× |
| c17 | 475.3 | 410.5 | +16% | 260.0 | 314.9 | −17% | 1.30× |
| dec2to4 | 422.4 | 346.9 | +22% | 260.0 | 276.4 | −6% | 1.26× |
| mux2 | 244.5 | 262.6 | −7% | 244.5 | 262.6 | −7% | 1.00× |
| and8 | 1424.1 | 576.7 | +147% | 610.6 | 639.9 | −5% | 0.90× |
| and8_tree | 217.8 | 249.2 | −13% | 217.8 | 249.2 | −13% | 1.00× |
The errors fall into three groups. The causes below are likely, not established.
Buffered netlists are 5% to 24% slower than predicted. At \(f \approx 3\) parasitic delay is a large share of each stage, and the model uses \(p_\text{inv} = 1\) while the calibration measures 3.51. Substituting 3.51 does not fix it, though: it overcorrects (for inv, \(3(4 + 3.51) = 22.5\,\tau = 218\,ps\) against 177 ps), because the fitted intercept also absorbs input slope effects that are smaller behind a fast, sized driver.
Single high-effort stages with series stacks are overestimated. The inverter, on which \(\tau\) is calibrated, is within 5%. NAND2, AOI21 and OAI22 are 22% to 51% faster than predicted. The model assumes a stack of \(n\) transistors is \(n\) times weaker, but in a short-channel process velocity saturation makes series stacks less than proportionally slower [7], so the textbook \(g\) is pessimistic for these gates.
and8 exposes a flaw in the sizing formulation. The importer turns the 8 input AND into a linear chain, so input \(a\) passes through 14 stages and input \(h\) through 2. The heaviest input is \(h\), next to the load, so the bisection sets \(f\) from its two stage path: \(F = 1.272 \times 64 = 81.4\) and \(f = \sqrt{81.4} = 9.02\). Applied to the long path, the same \(f\) shrinks every earlier gate by a factor of 9: from the output the drives are 7.1, 0.79, 0.11, 0.012, 0.002, and effectively zero for the first eight stages. Those gates are clamped up to the minimum width, where each drives a gate of its own size, so in silicon they are fast fanout of one stages rather than stages of effort 9. The model still charges \(14 \times 9.02 + P = 147\,\tau\), and the estimate is 147% high. The reported \(F = f^{14} = 2.4 \times 10^{13}\) is not a physical path effort, since the path’s input capacitance is effectively zero, and buffering built on it adds eight inverters that only add delay: 11% slower in SPICE. A single stage effort chosen from the input capacitance limit works when all paths have similar depth; it fails on unbalanced logic.
The benchmark and8_tree tests this directly. It is the
same function built as a balanced tree, \(((a
b)(c d))((e f)(g h))\), so every input is 6 stages deep. The path
effort is \(F = 1.272^3 \times 64 =
131.6\), every stage runs at \(f =
131.6^{1/6} = 2.26\), no gate is undersized, and no buffering is
needed since \(\log_4 131.6 = 3.5 <
6\). The estimate falls to within 13% of SPICE, inside the band
of the other benchmarks, and the measured delay is 249 ps against the
chain’s best of 577 ps: 2.3× faster with the same 42 transistors. The
error in and8 comes from the structure of the netlist, not from the
gates in it.
The full check, 22 macro_gen runs and 22 SPICE decks, takes about 3 minutes. Each macro_gen run spends about 50 s characterizing the inverter and under 1 s on import, mapping, sizing, buffering and output.
7. Summary
macro_gen now takes combinational logic in CIRCT, builds every gate as a static CMOS pull-up and pull-down network, sizes it by logical effort relative to one characterized inverter, and buffers each output to \(\operatorname{round}(\log_4 F)\) stages. It writes a transistor-level SPICE netlist, a structural Verilog netlist ready for physical design, and a report, and it logs every sizing step so each width and delay can be checked by hand. Once the inverter is characterized, sizing a design takes under a second.
Checked against SPICE on eleven SKY130 benchmarks, after calibrating \(\tau = 9.67\,\text{ps}\), \(p_\text{inv} = 3.51\) and \(C_\text{inv} = 1.93\,\text{fF}\) from simulation, the estimates fall 5% to 24% below the measured delay for every buffered netlist and within 24% for eight of eleven unbuffered ones. Buffering cuts the measured delay by up to 3.8×. Three lessons came out of the comparison: the textbook \(p_\text{inv} = 1\) and the FO4/5 rule do not hold in this process; the textbook \(g\) is pessimistic for series stacks; and a single stage effort for the whole netlist fails on unbalanced logic, as the 8 input AND shows, while the same function as a balanced tree lands within 13% and runs 2.3× faster. Feeding the calibrated constants back into the sizer, sizing along the critical path, and technology mapping are the next steps.