A copy of this work was available on the public web and has been preserved in the Wayback Machine. The capture dates from 2018; you can also visit <a rel="external noopener" href="https://www.pure.ed.ac.uk/ws/files/24271826/TACO_2014_1.pdf">the original URL</a>. The file type is <code>application/pdf</code>.
Efficient Power Gating of SIMD Accelerators Through Dynamic Selective Devectorization in an HW/SW Codesigned Environment
<span title="2014-07-31">2014</span>
<i title="Association for Computing Machinery (ACM)">
<a target="_blank" rel="noopener" href="https://fatcat.wiki/container/jfrn2kjyarhe7npmgvoxdp4cxu" style="color: black;">ACM Transactions on Architecture and Code Optimization (TACO)</a>
</i>
Leakage energy is a growing concern in current and future microprocessors. Functional units of microprocessors are responsible for a major fraction of this energy. Therefore, reducing functional unit leakage has received much attention in the recent years. Power gating is one of the most widely used techniques to minimize leakage energy. Power gating turns off the functional units during the idle periods to reduce the leakage. Therefore, the amount of leakage energy savings is directly
<span class="external-identifiers">
<a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1145/2629681">doi:10.1145/2629681</a>
<a target="_blank" rel="external noopener" href="https://fatcat.wiki/release/pidgahuxs5hhbfgfkzonpuhvta">fatcat:pidgahuxs5hhbfgfkzonpuhvta</a>
</span>
more »
... nal to the idle time duration. This paper focuses on increasing the idle interval for the higher SIMD lanes. The applications are profiled dynamically, in a Hardware/Software co-designed environment, to find the higher SIMD lanes usage pattern. If the higher lanes need to be turned-on for small time periods, the corresponding portion of the code is devectorized to keep the higher lanes off. The devectorized code is executed on the lowest SIMD lane. Our experimental results show that the average energy savings of the proposed mechanism are 15%, 12% and 71% greater than power gating, for SPECFP2006, Physicsbench and Eigen benchmark suites respectively. Moreover, the slowdown caused due to devectorization is negligible.
<a target="_blank" rel="noopener" href="https://web.archive.org/web/20180722095345/https://www.pure.ed.ac.uk/ws/files/24271826/TACO_2014_1.pdf" title="fulltext PDF download" data-goatcounter-click="serp-fulltext" data-goatcounter-title="serp-fulltext">
<button class="ui simple right pointing dropdown compact black labeled icon button serp-button">
<i class="icon ia-icon"></i>
Web Archive
[PDF]
<div class="menu fulltext-thumbnail">
<img src="https://blobs.fatcat.wiki/thumbnail/pdf/ac/c7/acc77488cc36fcb036c740104373aab9c2ef6adf.180px.jpg" alt="fulltext thumbnail" loading="lazy">
</div>
</button>
</a>
<a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1145/2629681">
<button class="ui left aligned compact blue labeled icon button serp-button">
<i class="external alternate icon"></i>
acm.org
</button>
</a>