{"id":1350,"date":"2013-11-03T10:59:26","date_gmt":"2013-11-03T09:59:26","guid":{"rendered":"http:\/\/htor.inf.ethz.ch\/blog\/?p=1350"},"modified":"2013-11-03T11:05:37","modified_gmt":"2013-11-03T10:05:37","slug":"advanced-mpi-programming-tutorial-at-supercomputing-2013","status":"publish","type":"post","link":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/2013\/11\/03\/advanced-mpi-programming-tutorial-at-supercomputing-2013\/","title":{"rendered":"Advanced MPI Programming Tutorial at Supercomputing 2013"},"content":{"rendered":"<p>Pavan Balaji, Jim Dinan, Rajeev Thakur and I are giving our <a href=\"http:\/\/sc13.supercomputing.org\/schedule\/event_detail.php?evid=tut128\" target=\"_blank\">Advanced MPI Programming tutorial<\/a> at Supercomputing 2013 on Sunday November 17th.<\/p>\n<p>Are you wondering about the new MPI-3 standard? How it affects you as a scientific or HPC programmer and what nice new features you can use to make your life easier and your application faster? Then you should not miss our tutorial.<\/p>\n<p><center><img src=\"http:\/\/www.unixer.de\/blog\/wp-content\/uploads\/mpi3v2.png\"><\/center><\/p>\n<p>Our abstract summarizes the main topics:<\/p>\n<blockquote><p>\nThe vast majority of production parallel scientific applications today use MPI and run successfully on the largest systems in the world. For example, several MPI applications are running at full scale on the Sequoia system (on ?1.6 million cores) and achieving 12 to 14 petaflops\/s of sustained performance. At the same time, the MPI standard itself is evolving (MPI-3 was released late last year) to address the needs and challenges of future extreme-scale platforms as well as applications. This tutorial will cover several advanced features of MPI, including new MPI-3 features, that can help users program modern systems effectively. Using code examples based on scenarios found in real applications, we will cover several topics including efficient ways of doing 2D and 3D stencil computation, derived datatypes, one-sided communication, hybrid (MPI + shared memory) programming, topologies and topology mapping, and neighborhood and nonblocking collectives. Attendees will leave the tutorial with an understanding of how to use these advanced features of MPI and guidelines on how they might perform on different<br \/>\nplatforms and architectures.\n<\/p><\/blockquote>\n<p>This tutorial is about advanced use of MPI. It will cover several advanced features that are part of<br \/>\nMPI-1 and MPI-2 (derived datatypes, one-sided communication, thread support, topologies and topology<br \/>\nmapping) as well as new features that were recently added to MPI as part of MPI-3 (substantial additions<br \/>\nto the one-sided communication interface, neighborhood collectives, nonblocking collectives, support for<br \/>\nshared-memory programming).<\/p>\n<p>Implementations of MPI-2 are widely available both from vendors and open-source projects. In addition,<br \/>\nthe latest release of the MPICH implementation of MPI supports all of MPI-3. Vendor implementations<br \/>\nderived from MPICH will soon support these new features. As a result, users will be able to use in practice<br \/>\nwhat they learn in this tutorial.<\/p>\n<p>The tutorial will be example driven, reflecting scenarios found in real applications. We will begin with<br \/>\na 2D stencil computation with a 1D decomposition to illustrate simple Isend\/Irecv based communication.<\/p>\n<p>We will then use a 2D decomposition to illustrate the need for MPI derived datatypes. We will introduce<br \/>\na simple performance model to demonstrate what performance can be expected and compare it with actual<br \/>\nperformance measured on real systems. This model will be used to discuss, evaluate, and motivate the rest<br \/>\nof the tutorial.<br \/>\nWe will use the same 2D stencil example to illustrate various ways of doing one-sided communication in<br \/>\nMPI and discuss the pros and cons of the different approaches as well as regular point-to-point communica-<br \/>\ntion. We will then discuss a 3D stencil without getting into complicated code details.<br \/>\nWe will use examples of distributed linked lists and distributed locks to illustrate some of the new ad-<br \/>\nvanced one-sided communication features, such as the atomic read-modify-write operations.<br \/>\nWe will discuss the support for threads and hybrid programming in MPI and provide two hybrid ver-<br \/>\nsions of the stencil example: MPI+OpenMP and MPI+MPI. The latter uses the new features in MPI-3 for<br \/>\nshared-memory programming. We will also discuss performance and correctness guidelines for hybrid pro-<br \/>\ngramming.<\/p>\n<p>We will introduce process topologies, topology mapping, and the new \u201cneighborhood\u201d collective func-<br \/>\ntions added in MPI-3. These collectives are particularly intended to support stencil computations in a scalable<br \/>\nmanner, both in terms of memory consumption and performance.<br \/>\nWe will conclude with a discussion of other features in MPI-3 not explicitly covered in this tutorial<br \/>\n(interface for tools, Fortran 2008 bindings, etc.) as well as a summary of recent activities of the MPI Forum<br \/>\nbeyond MPI-3.<\/p>\n<p>Our planned agenda for the day is<\/p>\n<ol>\n<li>Introduction (8.30\u201310.00)\n<ul>\n<li> Background: What is MPI\n<li> MPI-1, MPI-2, MPI-3\n<li> 2D stencil code with 1D decomposition: Isend\/Irecv version\n<li> 2D stencil code with 2D decomposition: Introduce derived datatypes\n<li> Introduce simple performance modeling and measurement\n<\/ul>\n<li> One-Sided Communication (10.30\u201312.00)\n<ul>\n<li> Basics of one-sided communication or remote memory access (RMA)\n<li> 2D stencil code with 1D decomposition: RMA with 3 forms of synchronization\n<li> 3D stencil: What changes and what to pay attention to\n<li> Introduce other features of MPI-3 RMA\n<li> Linked list or distributed lock example demonstrating new MPI-3 RMA features\n<\/ul>\n<li> Lunch (12.00\u20131.30)\n<li> MPI and Threads (1.30\u20133.00)\n<ul>\n<li> What does the MPI standard specify about threads\n<li> How does it enable hybrid programming\n<li> Hybrid (MPI+OpenMP) version of 2D stencil code\n<li> Hybrid (MPI+MPI) version of 2D stencil code using MPI-3 shared-memory support\n<li> Performance and correctness guidelines for hybrid programming\n<\/ul>\n<li> Topologies, Neighborhood\/Nonblocking Collectives (3.30-5.00)\n<ul>\n<li> Topologies and topology mapping\n<li> 2D stencil code with 2D decomposition using neighborhood collectives\n<li> MPI-3 nonblocking collectives with example\n<li> Summary of other features in MPI-3\n<li> Summary of recent activities of the MPI Forum\n<li> Conclusions\n<\/ul>\n<\/ol>\n<p>We&#8217;re looking forward to many interesting discussions!<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Pavan Balaji, Jim Dinan, Rajeev Thakur and I are giving our Advanced MPI Programming tutorial at Supercomputing 2013 on Sunday November 17th. Are you wondering about the new MPI-3 standard? How it affects you as a scientific or HPC programmer &hellip; <a class=\"more-link\" href=\"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/2013\/11\/03\/advanced-mpi-programming-tutorial-at-supercomputing-2013\/\">read more<\/a><\/p>","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":[],"categories":[5,2],"tags":[],"_links":{"self":[{"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/posts\/1350"}],"collection":[{"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/comments?post=1350"}],"version-history":[{"count":9,"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/posts\/1350\/revisions"}],"predecessor-version":[{"id":1359,"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/posts\/1350\/revisions\/1359"}],"wp:attachment":[{"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/media?parent=1350"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/categories?post=1350"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/htor.inf.ethz.ch\/blog\/index.php\/wp-json\/wp\/v2\/tags?post=1350"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}