macro_gen Dev Log #2: Sizing Combinational Logic by Logical Effort

macro_gen now sizes arbitrary combinational CMOS logic by logical effort, relative to the inverter characterized in part 1, and buffers outputs to log4(F) stages. Method with a fully worked example, calibration of the model constants, and a comparison against SPICE on eleven SKY130 benchmarks.

Summary. We taught a computer to size transistors for combinational logic. Given a design in the hw and comb dialects of CIRCT [1], [2] and a PDK, macro_gen (v0.2.0) builds every gate as a static CMOS pull-up and pull-down network, sizes each transistor by the method of logical effort [3] relative to the reference inverter characterized in part 1, and inserts inverters so each output path has the delay-optimal number of stages. We calibrated the model’s constants in ngspice [4] and compared its delay estimates against SPICE on eleven benchmarks in SKY130 [5]. For every buffered netlist the estimate is 5% to 24% below the measured delay. For unbuffered netlists the estimate is within 24% on eight of eleven designs and overestimates by 45% to 147% on the other three. In SPICE, buffering reduces the worst case delay by 1.3× to 3.8× on eight designs, leaves two unchanged, and makes one 11% slower. The worst case, an 8 input AND built as a linear chain, comes within 13% when the same function is built as a balanced tree, which is also 2.3× faster.

1. Problem

Part 1 sized one inverter to a target switching threshold by simulation: an 11 point ngspice sweep, about 50 s per design point. Repeating that for every gate type does not scale. Logical effort avoids it: once one inverter is characterized, every other gate can be sized analytically relative to it. This entry applies that method automatically to arbitrary combinational netlists and measures how well it predicts delay.

2. Model

The delay of a logic stage, in units of a process constant \(\tau\), is

\[ d = f + p, \qquad f = g\,h, \qquad h = \frac{C_\text{out}}{C_\text{in}}, \]

where \(g\) is the stage’s logical effort (its input capacitance relative to an inverter delivering the same current), \(h\) its electrical effort, and \(p\) its parasitic delay. Along a path of \(N\) stages, the path effort is

\[ F = G\,B\,H, \qquad G = \prod_i g_i, \qquad B = \prod_i b_i, \qquad H = \frac{C_\text{load}}{C_{\text{in},1}}, \]

where \(b_i = (C_\text{on path} + C_\text{off path}) / C_\text{on path}\) is the branching effort at stage \(i\). Delay is minimized when every stage carries the same effort, and the optimal number of stages is

\[ \hat f = F^{1/N}, \qquad \hat N = \operatorname{round}\left(\log_\rho F\right), \quad \rho \approx 4. \]

Transistor sizing. The reference inverter has \(W_n = 0.42\,\mu\text{m}\) and \(W_p = 1.1269\,\mu\text{m}\), so \(\gamma = W_p / W_n\). Each gate is sized so its worst pull-up and pull-down paths conduct like the inverter’s: a transistor in a series stack of \(n\) is \(n\) times wider. With widths in units of the inverter’s NMOS, the logical effort of input \(i\) is

\[ g_i = \frac{W_{n,i} + \gamma\,W_{p,i}}{1 + \gamma}. \]

Netlist sizing. Real netlists branch and reconverge, so instead of enumerating paths macro_gen gives every stage the same effort \(f\). In reverse topological order each stage gets drive \(s\) and input capacitance

\[ s = \frac{C_\text{out}}{f}, \qquad C_{\text{in},i} = g_i\, s, \]

where \(C_\text{out}\) is the sum of the input capacitances it drives plus \(C_\text{load}\) at an output port. The largest primary input capacitance falls monotonically as \(f\) rises, so \(f\) is found by bisection such that \(\max C_\text{in} = C_{\text{in},\max}\). On a single path this reproduces \(\hat f = F^{1/N}\).

Buffering. Each output whose path has fewer than \(\hat N\) stages gets \(\hat N - N\) inverters before its port, and the netlist is resized.

3. Using macro_gen

A run is described by one TOML file: the PDK, the reference inverter, the sizing targets and the design. The part that matters here is

[sizing]
cload_cinv = 64.0  # load on each module output port, in C_inv
cin_cinv = 1.0     # largest capacitance any primary input may present
stage_effort = 4.0 # rho, for buffering

[circt]
mlir_path = "c17.mlir"
top_module = "c17"

The design is plain CIRCT. c17 [6] has six NAND gates; since comb has no NAND, each is an AND followed by a NOT, written as comb.xor with a constant true:

hw.module @c17(in %N1: i1, in %N2: i1, in %N3: i1, in %N6: i1, in %N7: i1,
               out N22: i1, out N23: i1) {
  %true = hw.constant true
  %a10 = comb.and %N1, %N3 : i1
  %N10 = comb.xor %a10, %true : i1
  ...
  hw.output %N22, %N23 : i1, i1
}
macro_gen --config benchmarks/c17/c17.toml --emit-verilog --add-buffer

The run log shows every sizing decision, so any number in the netlist can be traced by hand. For each cell, in signal order: its logical effort and parasitic delay, the capacitance on each input and on its output, its drive, and the resulting transistor widths (* marks a width clamped up to the process minimum):

Sized cells (unbuffered): every stage at f = C_out / s = 14.385, C_in(pin) = g x s
  cell    | master      | g (per input) | p (tau) | C_in (C_inv)  | C_out (C_inv) | C_out (fF) | drive s (x inv) | NMOS W (um) | PMOS W (um)
  --------+-------------+---------------+---------+---------------+---------------+------------+-----------------+-------------+------------
  and2_2  | nand2_x0p10 | 1.272 / 1.272 | 2.000   | 0.133 / 0.133 | 1.500         | 2.23       | 0.104           | 0.42*x2     | 0.42*x2
  and2_6  | nand2_x0p39 | 1.272 / 1.272 | 2.000   | 0.500 / 0.500 | 5.657         | 8.43       | 0.393           | 0.42*x2     | 0.44x2
  and2_4  | nand2_x0p79 | 1.272 / 1.272 | 2.000   | 1.000 / 1.000 | 11.314        | 16.85      | 0.786           | 0.66x2      | 0.89x2
  and2_10 | nand2_x4p45 | 1.272 / 1.272 | 2.000   | 5.657 / 5.657 | 64.000        | 95.34      | 4.449           | 3.74x2      | 5.01x2
  and2_0  | nand2_x0p39 | 1.272 / 1.272 | 2.000   | 0.500 / 0.500 | 5.657         | 8.43       | 0.393           | 0.42*x2     | 0.44x2
  and2_8  | nand2_x4p45 | 1.272 / 1.272 | 2.000   | 5.657 / 5.657 | 64.000        | 95.34      | 4.449           | 3.74x2      | 5.01x2

For each output, the critical path and every term of the path effort:

Output paths (unbuffered): F = G x B x H, f = F^(1/N), N_opt = log_4 F, delay = N f + P
  output | critical path                        | N | G     | B     | H      | F         | f      | N_opt | added inv | P (tau) | delay (tau)
  -------+--------------------------------------+---+-------+-------+--------+-----------+--------+-------+-----------+---------+------------
  N22    | N6 > and2_2 > and2_4 > and2_8 > N22  | 3 | 2.056 | 3.000 | 482.72 | 2976.9555 | 14.385 | 5.77  | 3         | 6.00    | 49.16
  N23    | N6 > and2_2 > and2_6 > and2_10 > N23 | 3 | 2.056 | 3.000 | 482.72 | 2976.9555 | 14.385 | 5.77  | 3         | 6.00    | 49.16

The same tables are printed again after buffering, followed by a summary:

Summary: c17
  metric            | unbuffered | buffered
  ------------------+------------+---------
  stage effort f    | 14.386     | 2.982
  cells             | 6          | 12
  transistors       | 24         | 36
  worst delay (tau) | 49.16      | 26.89
  worst delay (FO4) | 9.83       | 5.38

The capacitances in fF use the tool’s own estimate of \(C_\text{inv}\) (1.49 fF, from the gate capacitance at one bias point); Section 5 shows the effective value is 1.93 fF.

Verilog. With --emit-verilog the sized, buffered design is written as a structural netlist: cell instances and wires only, no assign and no behavioural code, so it can go straight to physical design. Each cell master is named for its stage and drive, nand2_x0p81 is a NAND2 at drive 0.81, and the same names are used in the SPICE netlist so the two line up for LVS. The buffering is visible at the end: three inverters of growing size, 2.41, 7.20 and 21.46, in front of each output, and a header comment records that both outputs are now inverted.

// c17: logical-effort sized structural netlist
// generated by macro_gen 0.2.0 (864ae133bc) on 2026-09-29 17:02:59 UTC
// macro_gen.inverted.N22 = true
// macro_gen.inverted.N23 = true
module c17 (N1, N2, N3, N6, N7, N22, N23);
  input N1;
  input N2;
  input N3;
  input N6;
  input N7;
  output N22;
  output N23;
  wire N22_buf0_y;
  wire N22_buf1_y;
  wire N22_pre;
  wire N23_buf0_y;
  wire N23_buf1_y;
  wire N23_pre;
  wire net_1;
  wire net_3;
  wire net_5;
  wire net_7;

  nand2_x0p35 and2_0 (.in0(N1), .in1(N3), .y(net_1));
  nand2_x0p44 and2_2 (.in0(N3), .in1(N6), .y(net_3));
  nand2_x0p69 and2_4 (.in0(N2), .in1(net_3), .y(net_5));
  nand2_x0p35 and2_6 (.in0(net_3), .in1(N7), .y(net_7));
  nand2_x0p81 and2_8 (.in0(net_1), .in1(net_5), .y(N22_pre));
  nand2_x0p81 and2_10 (.in0(net_5), .in1(net_7), .y(N23_pre));
  inv_x2p41 N22_buf0 (.in0(N22_pre), .y(N22_buf0_y));
  inv_x7p20 N22_buf1 (.in0(N22_buf0_y), .y(N22_buf1_y));
  inv_x21p46 N22_buf2 (.in0(N22_buf1_y), .y(N22));
  inv_x2p41 N23_buf0 (.in0(N23_pre), .y(N23_buf0_y));
  inv_x7p20 N23_buf1 (.in0(N23_buf0_y), .y(N23_buf1_y));
  inv_x21p46 N23_buf2 (.in0(N23_buf1_y), .y(N23));
endmodule

SPICE. Every master becomes a transistor-level subcircuit. The NAND2 above at drive 0.81 has two series NMOS at \(2 \times 0.81 \times 0.42 = 0.68\,\mu\text{m}\) and two parallel PMOS at \(0.81 \times 1.1269 = 0.91\,\mu\text{m}\):

.subckt nand2_x0p81 in0 in1 y VDD GND
XMN0 y in0 nn0 GND sky130_fd_pr__nfet_01v8 L=0.1500 W=0.6804 nf=1 mult=1
XMN1 nn0 in1 GND GND sky130_fd_pr__nfet_01v8 L=0.1500 W=0.6804 nf=1 mult=1
XMP0 y in0 VDD VDD sky130_fd_pr__pfet_01v8 L=0.1500 W=0.9128 nf=1 mult=1
XMP1 y in1 VDD VDD sky130_fd_pr__pfet_01v8 L=0.1500 W=0.9128 nf=1 mult=1
.ends nand2_x0p81

Report. reports/c17.toml carries the same results in machine-readable form, with units in every key, for scripts and design space exploration loops:

[unbuffered]
stage_effort = 14.3855
cells = 6
transistors = 24
delay_tau = 49.1564
delay_fo4 = 9.8313

[buffered]
stage_effort = 2.9821
cells = 12
transistors = 36
delay_tau = 26.8924
delay_fo4 = 5.3785

[[outputs]]
name = "N22"
logic_stages = 3
path_effort = 2976.9556
added_inverters = 3
delay_tau = 49.1564
buffered_delay_tau = 26.8924

4. Worked example: ISCAS-85 c17

The steps below reproduce the tables in Section 3 by hand. c17 [6] is six NAND2 gates. Take its worst path, \(N_6 \rightarrow \text{NAND}_\text{a} \rightarrow \text{NAND}_\text{b} \rightarrow \text{NAND}_\text{c} \rightarrow N_{22}\) (cells and2_2, and2_4 and and2_8), with \(C_\text{load} = 64\,C_\text{inv}\) and \(C_{\text{in},\max} = 1\,C_\text{inv}\).

Step 1: process ratio. \[\gamma = \frac{1.1269}{0.42} = 2.683\]

Step 2: logical effort of a NAND2. Its two NMOS are in series (width 2 each) and its two PMOS in parallel (width 1 each, scaled by \(\gamma\)): \[g_\text{NAND2} = \frac{2 + \gamma}{1 + \gamma} = \frac{4.683}{3.683} = 1.272\]

Step 3: path logical effort. \[G = g_\text{NAND2}^3 = 1.272^3 = 2.056\]

Step 4: branching effort. From the sized netlist, \(\text{NAND}_\text{a}\)’s output drives \(\text{NAND}_\text{b}\) (1.000 \(C_\text{inv}\)) and another NAND off the path (0.500), and \(\text{NAND}_\text{b}\)’s output drives \(\text{NAND}_\text{c}\) and a second NAND of equal size (5.657 each): \[B = \frac{1.000 + 0.500}{1.000} \cdot \frac{5.657 + 5.657}{5.657} = 1.5 \times 2 = 3.0\]

Step 5: electrical effort. The path enters \(\text{NAND}_\text{a}\), whose input presents 0.1326 \(C_\text{inv}\): \[H = \frac{64}{0.1326} = 482.7\]

Step 6: path effort and stage effort. \[F = G\,B\,H = 2.056 \times 3.0 \times 482.7 = 2977, \qquad \hat f = 2977^{1/3} = 14.39\]

Step 7: size the last stage. \(\text{NAND}_\text{c}\) drives the load directly: \[s = \frac{64}{14.39} = 4.449\] \[W_n = 2 \times 4.449 \times 0.42 = 3.74\,\mu\text{m}, \qquad W_p = 1 \times 4.449 \times 1.1269 = 5.01\,\mu\text{m}\]

Step 8: delay. A NAND2’s parasitic delay is \(p = 2\) in this model, so \[d = N(\hat f + p) = 3\,(14.39 + 2) = 49.2\,\tau\]

Step 9: buffering. \[\hat N = \operatorname{round}(\log_4 2977) = \operatorname{round}(5.77) = 6\] Three inverters are added (\(p_\text{inv} = 1\) in the model). After resizing, every stage runs at \(f = 2.98\): \[d = 3\,(2.98 + 2) + 3\,(2.98 + 1) = 26.9\,\tau\]

Step 10: compare with SPICE. With \(\tau = 9.67\,\text{ps}\) from Section 5, the unbuffered estimate is \(49.2 \times 9.67 = 475\,\text{ps}\) against \(410.5\,\text{ps}\) measured (+16%), and the buffered estimate is \(26.9 \times 9.67 = 260\,\text{ps}\) against \(314.9\,\text{ps}\) measured (−17%).

5. Calibration

\(\tau\), \(p_\text{inv}\) and the physical size of \(C_\text{inv}\) were measured, not assumed. The reference inverter drove \(h\) copies of itself, from an input edge shaped by a fanout of four driver, and the average of the rising and falling delays was recorded:

Fanout 1 2 4 8 16 32 64
Delay (ps) 37.5 51.0 72.9 111.9 188.4 342.2 653.3

A least squares fit of \(d(h) = \tau\,(p_\text{inv} + h)\) over \(h \ge 4\) (smaller fanouts are distorted by the input edge) gives

\[ \tau = 9.67\,\text{ps}, \qquad p_\text{inv} = 3.51, \qquad d_\text{FO4} = 72.9\,\text{ps}. \]

A capacitor matching the \(h = 64\) delay is 123.4 fF, so the effective input capacitance is \(C_\text{inv} = 1.93\,\text{fF}\) and the 64 \(C_\text{inv}\) load is 123.4 fF.

Two earlier shortcuts were wrong. Taking \(\tau = d_\text{FO4}/5\) assumes \(p_\text{inv} = 1\) and gives 14.6 ps, 50% too high. Estimating \(C_\text{inv}\) from the gate capacitance at a single DC bias point gave 1.49 fF, 23% too low, likely because it misses the Miller contribution of the gate to drain overlap while the inverter switches.

6. Validation against SPICE

Each benchmark was sized with and without buffering, and every netlist was simulated in ngspice with the SKY130 tt models at 1.8 V. Each output drove 123.4 fF. Every input that can change an output was toggled with a 20 ps edge, with the other inputs held at every combination that lets the change through (found by evaluating the design’s logic). For each (input, side inputs, output) the delay is the average of the rising and falling 50% to 50% delays; the worst of these is the measured delay. The estimate is the model’s delay in \(\tau\) times 9.67 ps.

Structure of the sized netlists:

Benchmark Logic stages \(F\) (worst path) Inverters added Transistors
inv 1 64 2 2 → 6
nand2 1 81.4 2 4 → 8
aoi21 1 128 3 6 → 12
oai22 1 128 3 8 → 14
and_or_chain 2 128 2 8 → 12
maj3 2 442.5 2 14 → 18
c17 3 2977 6 24 → 36
dec2to4 3 2316 11 28 → 50
mux2 5 655 0 20 → 20
and8 14 \(2.4 \times 10^{13}\) 8 42 → 58
and8_tree 6 131.6 0 42 → 42

Worst case delay, estimated and measured (ps). Error is (estimate − SPICE) / SPICE; speedup is the measured unbuffered delay over the measured buffered delay.

Benchmark Unbuffered est. SPICE Error Buffered est. SPICE Error Speedup
inv 628.4 596.5 +5% 145.0 177.1 −18% 3.37×
nand2 806.1 659.6 +22% 164.4 193.8 −15% 3.40×
aoi21 1273.6 879.9 +45% 195.1 229.1 −15% 3.84×
oai22 1276.2 847.2 +51% 197.8 259.8 −24% 3.26×
and_or_chain 264.5 299.5 −12% 195.1 229.1 −15% 1.31×
maj3 488.5 469.6 +4% 278.5 321.0 −13% 1.46×
c17 475.3 410.5 +16% 260.0 314.9 −17% 1.30×
dec2to4 422.4 346.9 +22% 260.0 276.4 −6% 1.26×
mux2 244.5 262.6 −7% 244.5 262.6 −7% 1.00×
and8 1424.1 576.7 +147% 610.6 639.9 −5% 0.90×
and8_tree 217.8 249.2 −13% 217.8 249.2 −13% 1.00×

The errors fall into three groups. The causes below are likely, not established.

Buffered netlists are 5% to 24% slower than predicted. At \(f \approx 3\) parasitic delay is a large share of each stage, and the model uses \(p_\text{inv} = 1\) while the calibration measures 3.51. Substituting 3.51 does not fix it, though: it overcorrects (for inv, \(3(4 + 3.51) = 22.5\,\tau = 218\,ps\) against 177 ps), because the fitted intercept also absorbs input slope effects that are smaller behind a fast, sized driver.

Single high-effort stages with series stacks are overestimated. The inverter, on which \(\tau\) is calibrated, is within 5%. NAND2, AOI21 and OAI22 are 22% to 51% faster than predicted. The model assumes a stack of \(n\) transistors is \(n\) times weaker, but in a short-channel process velocity saturation makes series stacks less than proportionally slower [7], so the textbook \(g\) is pessimistic for these gates.

and8 exposes a flaw in the sizing formulation. The importer turns the 8 input AND into a linear chain, so input \(a\) passes through 14 stages and input \(h\) through 2. The heaviest input is \(h\), next to the load, so the bisection sets \(f\) from its two stage path: \(F = 1.272 \times 64 = 81.4\) and \(f = \sqrt{81.4} = 9.02\). Applied to the long path, the same \(f\) shrinks every earlier gate by a factor of 9: from the output the drives are 7.1, 0.79, 0.11, 0.012, 0.002, and effectively zero for the first eight stages. Those gates are clamped up to the minimum width, where each drives a gate of its own size, so in silicon they are fast fanout of one stages rather than stages of effort 9. The model still charges \(14 \times 9.02 + P = 147\,\tau\), and the estimate is 147% high. The reported \(F = f^{14} = 2.4 \times 10^{13}\) is not a physical path effort, since the path’s input capacitance is effectively zero, and buffering built on it adds eight inverters that only add delay: 11% slower in SPICE. A single stage effort chosen from the input capacitance limit works when all paths have similar depth; it fails on unbalanced logic.

The benchmark and8_tree tests this directly. It is the same function built as a balanced tree, \(((a b)(c d))((e f)(g h))\), so every input is 6 stages deep. The path effort is \(F = 1.272^3 \times 64 = 131.6\), every stage runs at \(f = 131.6^{1/6} = 2.26\), no gate is undersized, and no buffering is needed since \(\log_4 131.6 = 3.5 < 6\). The estimate falls to within 13% of SPICE, inside the band of the other benchmarks, and the measured delay is 249 ps against the chain’s best of 577 ps: 2.3× faster with the same 42 transistors. The error in and8 comes from the structure of the netlist, not from the gates in it.

The full check, 22 macro_gen runs and 22 SPICE decks, takes about 3 minutes. Each macro_gen run spends about 50 s characterizing the inverter and under 1 s on import, mapping, sizing, buffering and output.

Reproduction bundle
macro-gen-dev-log-02-benchmarks.tar.gz
11 benchmark designs (CIRCT MLIR + config each), the benchmark runner, the SPICE check script and its results, and a README with exact repro steps · 10.5 KB

7. Summary

macro_gen now takes combinational logic in CIRCT, builds every gate as a static CMOS pull-up and pull-down network, sizes it by logical effort relative to one characterized inverter, and buffers each output to \(\operatorname{round}(\log_4 F)\) stages. It writes a transistor-level SPICE netlist, a structural Verilog netlist ready for physical design, and a report, and it logs every sizing step so each width and delay can be checked by hand. Once the inverter is characterized, sizing a design takes under a second.

Checked against SPICE on eleven SKY130 benchmarks, after calibrating \(\tau = 9.67\,\text{ps}\), \(p_\text{inv} = 3.51\) and \(C_\text{inv} = 1.93\,\text{fF}\) from simulation, the estimates fall 5% to 24% below the measured delay for every buffered netlist and within 24% for eight of eleven unbuffered ones. Buffering cuts the measured delay by up to 3.8×. Three lessons came out of the comparison: the textbook \(p_\text{inv} = 1\) and the FO4/5 rule do not hold in this process; the textbook \(g\) is pessimistic for series stacks; and a single stage effort for the whole netlist fails on unbalanced logic, as the 8 input AND shows, while the same function as a balanced tree lands within 13% and runs 2.3× faster. Feeding the calibrated constants back into the sizer, sizing along the critical path, and technology mapping are the next steps.

References

[1]
The CIRCT Authors, “CIRCT: Circuit IR compilers and tools.” 2026. Accessed: Sep. 29, 2026. [Online]. Available: https://circt.llvm.org
[2]
C. Lattner et al., “MLIR: Scaling compiler infrastructure for domain specific computation,” in 2021 IEEE/ACM international symposium on code generation and optimization (CGO), 2021, pp. 2–14. doi: 10.1109/CGO51591.2021.9370308.
[3]
I. Sutherland, B. Sproull, and D. Harris, Logical effort: Designing fast CMOS circuits. San Francisco, CA: Morgan Kaufmann, 1999.
[4]
The ngspice project, “Ngspice: Open source mixed-signal circuit simulator.” 2026. Accessed: Sep. 29, 2026. [Online]. Available: https://ngspice.sourceforge.io
[5]
SkyWater Technology Foundry and Google, “SkyWater SKY130 open source process design kit.” 2020. Accessed: Sep. 29, 2026. [Online]. Available: https://github.com/google/skywater-pdk
[6]
F. Brglez and H. Fujiwara, “A neutral netlist of 10 combinational benchmark circuits and a target translator in Fortran,” in Proceedings of the IEEE international symposium on circuits and systems (ISCAS), 1985, pp. 695–698.
[7]
N. H. E. Weste and D. M. Harris, CMOS VLSI design: A circuits and systems perspective, 4th ed. Boston, MA: Addison-Wesley, 2010.