摘要
High Bandwidth Memory couples many DRAM dies to a host through independent channels, and a memory controller presents each channel to the workload as either available or isolated. This paper takes that binary service interface as the modeling primitive and builds a closed-form framework for evaluating stack and system behavior on top of it. Each die is treated as a binary component whose reliability is composed from DRAM, through-silicon via and micro-bump contributions, with the via bundle itself modeled as a threshold subsystem; the dies are then aggregated as a threshold structure over the stack, and stacks are aggregated over the system. The result is an evaluation whose cost grows linearly rather than exponentially with the number of dies, and which yields not only reliability but the full distribution of delivered bandwidth, its moments, and the sensitivity of system availability to each component. Three design questions are answered directly: where to direct reliability investment, how many stacks to provision for a given availability target, and which bandwidth threshold minimizes cost when bandwidth and reliability requirements are imposed together. The approximations the framework makes are bounded rather than assumed. A three-state baseline quantifies the error introduced by the binary representation and shows it is governed by a single measurable quantity, and distribution-free inequalities bound the effect of correlation among component failure mechanisms, which proves negligible in the regime where HBM parts are qualified. Application to a representative stack identifies DRAM cell reliability as the dominant bottleneck and shows that a single spare via per bundle is sufficient at typical defect rates.