Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm
| dc.creator | Kowalkowski, Jim | |
| dc.date | 2003-06-13 | |
| dc.date.accessioned | 2026-07-07T11:46:13Z | |
| dc.date.available | 2026-07-07T11:46:13Z | |
| dc.description | When thousands of processors are involved in performing event filtering on a trigger farm, there is likely to be a large number of failures within the software and hardware systems. BTeV, a proton/antiproton collider experiment at Fermi National Accelerator Laboratory, has designed a trigger, which includes several thousand processors. If fault conditions are not given proper treatment, it is conceivable that this trigger system will experience failures at a high enough rate to have a negative impact on its effectiveness. The RTES (Real Time Embedded Systems) collaboration is a group of physicists, engineers, and computer scientists working to address the problem of reliability in large-scale clusters with real-time constraints such as this. Resulting infrastructure must be highly scalable, verifiable, extensible by users, and dynamically changeable. | |
| dc.description | Paper for the 2003 Computing in High Energy and Nuclear Physics (CHEP03), La Jolla, Ca, USA, March 2003. PSN THGT001 | |
| dc.identifier | https://arxiv.org/abs/cs/0306074 | |
| dc.identifier | http://arxiv.org/abs/cs/0306074 | |
| dc.identifier | ECONFC0303241:THGT001,2003 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/202149 | |
| dc.subject | Distributed, Parallel, and Cluster Computing | |
| dc.subject | B.8.1 | |
| dc.title | Understanding and Coping with Hardware and Software Failures in a Very Large Trigger Farm | |
| dc.type | text |