Full citation: M. Choi and L. Milor, “Diagnosis of Optical Lithography Faults With Product Test Sets,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 27, no. 9, pp. 1657–1669, Sept. 2008, doi: 10.1109/TCAD.2008.927672.
Plain-Language Overview
Modern integrated circuits do not manufacture perfectly uniformly across the surface of a chip. Optical lithography—the process used to transfer circuit patterns onto silicon—can systematically make transistor features slightly larger or smaller depending on nearby shapes, their position on the chip, and the surrounding pattern density. Those dimensional changes can alter transistor and circuit delays.
Choi and Milor propose diagnosing these lithography-related variations using the chip’s own product test patterns rather than relying only on dedicated process-monitor structures. Their central idea is that different lithographic mechanisms produce different changes in the timing of critical paths. Those changes can therefore create recognizable pass/fail signatures when carefully selected timing tests are applied.
The method combines physical-layout information, timing analysis, critical-path enumeration, test-pattern generation, lithography fault simulation, and correlation-based diagnosis. The resulting diagnostic system attempts to identify the physical process mechanism responsible for abnormal within-die timing variation.
What Problem the Paper Addresses
Traditional process monitoring often uses test structures located in scribe lines around a die. This works reasonably well for variation that is strongly correlated across an entire chip or wafer. It is much less effective for systematic within-die variation, because lithographic effects can depend strongly on where a transistor is located and what geometries surround it.
The authors focus on three lithography-related sources of critical-dimension variation: the optical proximity effect, lens aberrations, and flare. These phenomena can change gate dimensions and consequently alter path delays.
The practical problem is therefore not merely determining that a chip has a timing problem. Manufacturing engineers need to identify which physical mechanism caused it, because different mechanisms call for different corrective actions. The paper notes, for example, that proximity effects can motivate local mask correction, spatial mask correction can address lens-related variation, and dummy-feature strategies can reduce density-related flare effects. Because these corrections increase mask cost, diagnosis can help target corrective effort rather than applying every correction indiscriminately.
Questions the Paper Answers
The paper examines whether product timing tests can reveal systematic within-die variation caused by lithography; whether critical and near-critical circuit paths can be efficiently extracted while accounting for input transition time; how layout-dependent lithographic effects can be incorporated into timing and fault simulation; whether different physical lithography mechanisms generate sufficiently different pass/fail signatures for diagnosis; and how fault detectability and diagnostic accuracy depend on the number of available test vectors and the magnitude of within-die variation.
Key Technical Terms and Definitions
| Term | Accessible definition |
|---|---|
| Within-die variation | Differences in device or interconnect characteristics at different locations within the same chip. |
| Systematic variation | Variation that follows a repeatable physical or spatial pattern rather than occurring randomly. |
| Critical dimension (CD) | A key manufactured feature dimension, such as transistor gate length, whose variation can affect performance. |
| Optical proximity effect | Lithographic variation caused by neighboring mask features changing the exposure received by a feature. |
| Coma | A lens aberration that can make printed dimensions depend asymmetrically on features located on different sides of a pattern. |
| Lens aberration | Imperfection in the optical projection system that makes printed dimensions depend on position in the exposure field. |
| Flare | Unwanted scattered or reflected light that can cause printed dimensions to depend on local mask-pattern density. |
| Path delay fault | A timing fault in which the propagation delay along a circuit path violates an allowed timing limit. |
| Critical/near-critical path | A circuit path whose delay is at or close to the maximum delay of the circuit. |
| Dynamic timing analysis | Timing analysis that propagates particular input transition sequences through a circuit rather than considering only worst-case static timing quantities. |
| Fault dictionary | A database associating modeled physical faults with the pass/fail signatures predicted for a set of test patterns. |
| Fault signature | The vector of passing and failing test outcomes produced by a particular fault. |
Workflow
The methodology can be summarized as follows:
- Model lithography-dependent variation. Transistor gate dimensions are altered according to physical neighborhood, location in the reticle, or pattern density. Neighborhood effects represent proximity and Coma behavior; spatial variation represents lens aberrations; density-dependent variation represents flare.
- Perform layout-dependent timing characterization. Layout information is attached to gate instances, transistor dimensions are modified according to the selected physical fault model, and affected cells are recharacterized. The implementation described in the paper uses HSPICE for gate characterization.
- Enumerate critical and near-critical paths. A depth-first-search algorithm is combined with backward signal propagation so that input transition time is included in path-delay estimates. Search branches that cannot exceed the timing threshold are pruned.
- Generate and compact test patterns. Sensitizable long paths are converted into product test vectors with a commercial automatic test-pattern-generation flow.
- Simulate each physical fault dynamically. The selected patterns are propagated using the modified timing information, producing path delays for every fault/test combination.
- Convert timing behavior into pass/fail signatures. Most experiments select paths exceeding roughly 90% of the maximum fault-free delay and use a test frequency based on \(1/(0.9d_{max})\). Systematic within-die variation can reorder path delays, causing some patterns to pass when the corresponding fault-free test population would fail.
- Construct the fault dictionary and diagnose silicon behavior. Simulated pass/fail vectors are stored by physical fault type and variation magnitude. An observed test vector is correlated with these dictionary entries, and high correlation with a particular physical-origin class supplies evidence for that mechanism.
Main Findings
The experiments use ISCAS’85 benchmark circuits. One important computational result is that pruning significantly reduces the cost of critical-path enumeration. The authors report that the pruned DFS implementation was about ten times faster on average than ordinary DFS, although circuits containing unusually large fractions of long paths, such as c432 and c2670, benefited less from pruning.
Fault detectability improves as the magnitude of systematic variation increases. The results summarized in Figure 15 on page 1667 show that a 15% range of variation is more readily detected than a 5% range. Figure 16 also shows an important relationship between detectability and test-set size: circuits for which many useful test vectors were generated tended to exhibit better detection performance.
In particular, c1908 and c5315 each produced more than 900 test vectors and showed the strongest detectability. The experiments further indicate that most tested lithography faults could be correctly distinguished for these circuits. In contrast, c7552 had fewer diagnostic vectors and poorer discrimination between fault mechanisms. The authors’ discussion ultimately characterizes diagnosis as successful when roughly 1000 or more suitable test vectors can be generated.
Figures 17–19 on page 1668 evaluate diagnosis when the actual variation does not necessarily equal one of the exact magnitudes represented in the fault dictionary. Faults of 5%, 7.5%, 10%, 12.5%, and 15% were tested against dictionaries containing different subsets of 5%, 10%, and 15% models. The correlation approach therefore does not require an exact simulated match: a fault can still be identified when its pass/fail pattern is most similar to dictionary entries belonging to the correct physical mechanism.
An important detail is that the experimental dictionary does not exercise every mechanism described in the modeling section. Table V uses eight physical-origin categories associated with proximity/Coma and directional lens-aberration behavior; the paper’s flare model is presented as part of the general methodology but is not included in that reported eight-origin experimental fault set.
Technical Significance
The paper moves delay-fault diagnosis from simply asking “Which path failed?” toward asking “Which manufacturing mechanism produced the timing distribution that caused the failure?”
That distinction is technically important. A systematic lithographic disturbance is distributed across many devices, so it does not behave like a single localized open, short, or defective gate. The use of a path-delay model therefore reflects the cumulative nature of the variation.
Another important contribution is the integration of layout context into timing analysis. Conventional cell timing models primarily describe circuit topology and technology-library behavior. The proposed framework additionally carries information describing transistor neighborhood, physical location, and pattern density, allowing a process mechanism to be propagated into circuit-level timing predictions.
The path-enumeration technique is also significant because it accounts for transition slope. The authors emphasize that the path with the latest-arriving input is not automatically the slowest path: a different input with a slower transition can produce a greater gate delay. Their backward-propagation tables retain maximum delay-to-output information as a function of transition time, making pruning more physically realistic than algorithms based on fixed edge delays.
Industrial Impact
The most direct manufacturing implication is the possibility of using product circuitry itself as a process diagnostic sensor. Dedicated scribe-line structures provide limited spatial coverage, while a functional design contains large numbers of transistors distributed throughout the reticle and across many different layout environments.
If the diagnostic signature can reliably identify a physical origin, engineers can potentially direct expensive mask or process corrections toward the mechanism actually affecting yield. The authors specifically connect diagnosis with decisions concerning optical proximity correction, spatial mask correction for lens effects, and approaches intended to make mask density more uniform.
The proposed method also provides a conceptual bridge between design-for-test, timing analysis, process characterization, and yield learning. Rather than treating electrical test results and lithography characterization as separate activities, the framework links an observed electrical signature back to a hypothesized process mechanism.
The authors further suggest averaging test results across multiple chips because systematic within-die lithographic variation is generally associated with a fabrication condition affecting many devices manufactured at approximately the same time. Such averaging could strengthen systematic signatures relative to random chip-specific defects.
Why the Paper Matters
As semiconductor dimensions shrink, variations that were once small compared with design margins become increasingly important. A chip can therefore satisfy functional logic testing while still experiencing timing-yield loss caused by systematic manufacturing variation.
This paper is notable because it treats timing-test data as more than a simple pass/fail quality gate. The same measurements can potentially contain information about why a product is behaving abnormally.
That concept remains broadly relevant: if physical manufacturing mechanisms can be modeled well enough to predict distinct electrical signatures, ordinary product measurements can become a form of indirect process metrology.
Limitations and Scope
The study is deliberately scoped to a restricted diagnostic problem. It focuses primarily on systematic gate-layer variation associated with optical lithography and evaluates the methodology using ISCAS’85 benchmark circuits and simulation, rather than reporting validation on production silicon.
Although flare is modeled conceptually, the principal reported experimental fault dictionary contains eight physical-origin categories based on optical proximity/Coma behavior and four directional lens-aberration cases.
The implemented dynamic timing analysis also omits several effects. The authors explicitly state that their implementation does not account for interconnect crosstalk, data-dependent delays, or the closeness dependency of multiple input transitions.
Diagnostic capability depends heavily on the availability of a sufficiently large and diverse test set. The poorer results for c7552 illustrate that a circuit with too few useful test vectors may not provide enough independent signatures to separate physical mechanisms.
The dictionary is necessarily a discretization of continuous process variation. Correlation mitigates the requirement for an exact magnitude match, but diagnosis still depends on whether the simulated dictionary contains models representative of the real physical mechanisms.
Finally, extending the framework to additional causes of within-die variation—such as supply-voltage variation, temperature variation, etch or CMP effects, and interconnect lithography—would expand the fault dictionary and probably increase the number of test patterns needed for reliable discrimination. The authors explicitly identify this dictionary growth as an important challenge for broader application.
Concise Technical Abstract
Choi and Milor present a product-test-based methodology for diagnosing systematic within-die timing variation caused by optical lithography. Physical fault models relate gate critical-dimension changes to layout neighborhood, reticle location, and pattern density. A layout-dependent timing framework recharacterizes affected cells, while transition-aware backward timing propagation and pruned depth-first search identify critical and near-critical paths. ATPG-derived patterns are then evaluated with dynamic timing analysis under modeled lithography faults. Each physical fault produces a pass/fail signature that is stored in a fault dictionary, and observed signatures are diagnosed through correlation. Experiments on ISCAS’85 benchmarks show that pruning substantially accelerates path enumeration, fault detectability rises with variation magnitude and test-set size, and lithographic fault mechanisms can generally be distinguished when sufficiently large diagnostic test sets—on the order of 1000 vectors in the authors’ experiments—are available.
Leave a comment