• DocumentCode
    2405243
  • Title

    A performance monitor based on virtual global time for clusters of PCs

  • Author

    Taufer, Michela ; Stricker, Thomas

  • Author_Institution
    Dept. of CSE, Califonia Univ., San Diego, CA, USA
  • fYear
    2003
  • fDate
    1-4 Dec. 2003
  • Firstpage
    64
  • Lastpage
    72
  • Abstract
    Debugging the performance of parallel and distributed systems remains a difficult task despite the widespread use of middleware packages for automatic distribution, communication and tasking in clusters. In this paper we present a performance monitoring tool for clusters of PCs that is based on the simple concept of accounting for resource usage and on the simple idea of mapping all performance related state of hardware performance counters and operating system variables backwards to the application level. In this way a monitoring tool can explain the most relevant performance metrics at a higher level that is easily understood by the application developer. The most important metric for distributed high performance applications remains the total execution time vs. the number of compute nodes involved, since it translates into the scalability of an application. As a detailed contribution of this paper, we closely look into what is needed to reverse map the low level performance counters at each node back through the middleware layer responsible for the parallelization and distribution. The specific problems encountered and dealt with are the creation of a flexible notion of global time for time-stamping and the reassembling of performance data and an appropriate communication mechanism to minimize monitoring intrusion due to the additional networking traffic caused by the monitor. We show how our tool can be used to measure, explain and predict the performance and scalability of a distributed OLAP application running on clusters of PCs.
  • Keywords
    parallel processing; performance evaluation; system monitoring; transaction processing; workstation clusters; OLAP application; PC clusters; automatic distribution; distributed systems; hardware performance counters; middleware; monitoring intrusion; networking traffic; operating system variables; parallel systems; performance analysis; performance evaluation; performance metrics; performance monitoring tool; program debugging; resource usage; time-stamping; traffic monitoring; virtual global time; Counting circuits; Debugging; Hardware; Measurement; Middleware; Monitoring; Operating systems; Packaging; Parallel processing; Personal communication networks; Scalability;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Cluster Computing, 2003. Proceedings. 2003 IEEE International Conference on
  • Print_ISBN
    0-7695-2066-9
  • Type

    conf

  • DOI
    10.1109/CLUSTR.2003.1253300
  • Filename
    1253300