The Role of Low Latency File Access in Accelerating AI Workloads

By marketing@site-a.com

The use of artificial intelligence (AI) is rapidly moving from the lab into the mainstream. The reason? Businesses believe AI can deliver operational cost savings, improve decision making, enhance customer interactions, speed data mining, and boost data security. As such, the number of companies using AI has grown by 270% in the past four years. As a result, organizations need to design high-performance computing architectures for AI workloads. 

Supporting AI efforts requires high-performance computing (HPC) capabilities to perform rapid analysis, tune neural net models, and conduct machine learning by examining large datasets. Fortunately, HPC requirements for AI are similar to other compute-intensive applications (e.g., Big Data analytics, forecasting, modeling, and finite element simulations) that are also increasingly being introduced into the enterprise today. That means there are many high-performance core compute, storage, and networking technologies available, which have made their way from supercomputing centers and academic labs into the enterprise.

However, several factors determine what type of infrastructure elements are needed for specific AI applications. Many AI efforts need speedy execution. That is the case for AI applications that do things like power autonomous systems, engage customers in real-time via chat or natural language, or identify outliners to prevent fraud. Such applications must run their analyses and get actionable information in real-time.

Selecting the right technology solution

To achieve the necessary performance, AI deployments typically make use of expensive GPU processing arrays. For workloads to run cost-effectively, there is a need to support high data rates to keep the processors satiated. That, in turn, dictates the use of ultrafast interconnect technology and tightly coupled high-performance storage.

Looking deeper, some of the GPU options include:

  • NVIDIA Tesla P100 GPU accelerators for PCIe based servers. Tesla P100 with NVIDIA NVLink delivers up to a 50X performance boost for the top HPC applications and all deep learning frameworks.
  • NVIDIA Tesla V100 Tensor Core, powered by NVIDIA Volta architecture, is a data center GPU to accelerate HPC and AI workloads.
  • NVIDIA T4 GPU, which accelerates cloud workloads, is used for HPC, deep learning training, and inference, machine learning, and data analytics.
  • GEFORCE RTX 2080 Ti is NVIDIA’s flagship graphics card based on NVIDIA Turing™ GPU architecture and ultra-fast GDDR6 memory.

Systems that use these and other GPUs to accelerate AI workloads need high-performance interconnect technologies to make cost-effective use of their performance capabilities. The internet technologies of choice include InfiniBand, Omni-Path, and remote direct memory access (RDMA). 

What are their capabilities?

  • InfiniBand is a computer-networking communications standard used in HPC systems that features very high throughput and very low latency. It is used as either a direct or switched interconnect between servers and storage systems, as well as an interconnect between storage systems.
  • Omni-Path(also Omni-Path Architecture or OPA) is a high-performance communication architecture from Intel. It delivers low communication latency and high throughput.
  • RDMA is an industry-standard that supports what is known as zero-copy networking by enabling the network adapter to move data directly to or from the application. This eliminates both the operating system and CPU involvement, so it is exceptionally faster than other solutions.

The plethora of GPU and interconnect technologies choices is a double-edged sword. On the plus side, the right combination will produce an optimized system to accelerate a specific AI application. On the downside, many businesses do not have expertise in these technologies and need help selecting the best solution for their application and optimizing a system’s performance.

When selecting elements in a system, it’s critical to determine which interconnect solution provides low latency file access to help AI and HPC workloads achieve higher performance and scalability.

These challenges exist when trying to configure a system for any HPC application, but the issues are especially important with AI applications. They need fast access to data to reduce training time in a deep learning scenario but also in supporting fast decision making in production environments.

Determining which is best for you

So how do you determine which storage is best for an AI application? Beyond basics like determining cost/performance issues when using hard drives versus solid-state and flash drives, there are storage file system and architecture issues to consider. Do you use a distributed architecture? Do you need a parallel file system? The bottom-line: AI applications need storage solutions that offer the highest throughput, lowest latency data access for CPU- and GPU-intensive AI, and HPC workloads.

In the final analysis, to optimize the running of AI workloads and make the most efficient use of expensive GPU arrays, compute solutions must bring together the right GPU, high-performance storage, and interconnect technologies. These technologies must be tightly integrated and tuned to optimize the solution’s performance when running AI workloads.

Top Tips to Improving Your System Administration Skills

By marketing@site-a.com

You Envision the Future

LAKE FOREST, CALIF., SEPT 12, 2019 — PSSC Labs, a developer of custom High Performance Computing and Big Data computing solutions today shares their top tips for improving your system administration skills.

At PSSC Labs, our mission is to deliver reliable, high performance turnkey platforms that help organizations like yours perform necessary data tasks without giving up control of your data.  That means keeping your system administration skills up-to-date is more important than ever! Of course, we’re always here to help you should you ever encounter any issues. In fact, you can contact our Support line at any time and we will be happy to assist you. But just to give you a little boost, we’ve outlined the top five support issues that on-premise model users encounter, and what to do if it happens to you.

Machine Not Powering On

PSSC Labs produces turnkey platforms, so they’re ready to be plugged in and get going upon delivery. Should you ever notice that one of your machines isn’t powering on, the first thing you should check is power supply. Verify that the power supply is receiving power from the source by verifying that the power connectors are securly plugged in on both ends.

If this doesn’t work, try a new power port on the PDF and then move on to a new power cable. If your machine still does not power on you can try using a power supply from your spare parts kit (if a spare parts kit was provided with your order) or from another working machine.  Machines with faulty memory can also exhibit this type of behavior, so shuffling or re-seating the memory is also worth trying. 

Machine Has Crashed or Rebooted

On-premise models can crash or reboot for various reasons. They’re often accompanied by BMC System Event Log messages.  Start by checking the output of the command ‘ipmitool sel list’ and check “/var/log/messages” for errors correlating to the date / time of the crash.

Updating your machine’s BIOS to the latest version is recommended as well, as it can fix crashing systems as well as improve the machine’s overall performance.  Verify that the CPUs are not overheating by checking the “sensors” command while the machine is under load.  Faulty memory can also cause crashes, so re-seating or shuffling memory can help to resolve these issues as well.  The causes of system crashes are sometimes hard to pin down, so do not hesitate to contact the PSSC Labs support team for troubleshooting assistance.

Software Support

PSSC Labs solutions often come pre-installed with multiple software applications, ranging from job submission software to parallel computing platforms, from weather modeling to data analysis.

Many of these packages are third party applications and not created by PSSC Labs, so if you ever have an issue with pre-loaded software, it’s recommended to first contact the software developers. If you still need assistance afterwards, contact the PSSC Labs support staff who can promptly assist you further.

Machine is Reporting Hard Drive/RAID Errors

This error is sometimes challenging for even those with years of experience and expert administration skills. Whenever your hard drive fails, your server’s data is at risk.  If you see errors indicating a possible issue with one of your hard drives, a PSSC Labs support technician can assist you with checking the drive’s built in SMART (Self-Monitoring, Analysis and Reporting Technology) data, which will show the drive’s health and indicate whether or not it needs to be replaced.  It’s possible that the drive is healthy and there is another underlying issue which needs to be remedied, so please contact PSSC Labs should you encounter this error.

Machine is Reporting Memory/CPU Errors

PSSC Labs servers come equipped with large amounts of memory and speedy CPUs. When issues occur with these components, users will often see slow / degraded system performance accompanied by MCE (Machine Check Error) messages.  You can contact PSSC Labs and provide these messages to a PSSC Labs support technician who can assist you with determining which memory DIMM or CPU the error is coming from.  Memory can be easily replaced, whereas handling CPUs is an extremely delicate process — one that we recommend avoiding unless absolutely necessary. Often times MCE messages can be resolved by simply reseating or shuffling the memory around.

About PSSC Labs 

For technology powered visionaries with a passion for challenging the status quo, PSSC Labs is the answer for hand-crafted HPC and Big Data computing solutions that deliver relentless performance with the absolute lowest total cost of ownership.  

 We are true innovators offering high performance computing solutions to solve the world’s most demanding problems. For 25+ years, organizations of all sizes and from a variety of sectors rely on PSSC Labs’ computing systems. We are proud to support many departments within the United States government, Fortune 500 companies, as well as small and medium-sized businesses.  

All products are designed and built at the company’s headquarters in Lake Forest, California.

Government Agencies Leverage New Technology for Data Flow

By marketing@site-a.com

LAKE FOREST, CALIF., AUGUST 20, 2019 — PSSC Labs, a developer of custom High Performance Computing and Big Data computing solutions today shares how the CyberRax Data Flow Pipleline is helping various DOD organizations achieve their goals.

The Federal Government has spent decades collecting and storing huge data sets, covering everything from census and population records to crime reports and infrastructure data. The analytics retrieved from this data has assisted our government agencies in commissioning greater responsiveness and efficiencies at every level, making it truly invaluable to our society. Federal agencies simply can’t afford to suffer from computational roadblocks. Government servers need to be able to perform streaming analytics while also providing full control over their own data, especially when it comes to governance, security and performance.   

PSSC Labs, in partnership with Cloudera, is deploying the CyberRax Data Flow Pipeline as a truly turnkey system engineered specifically to meet the data transport needs and security requirements of Federal Agencies. The CyberRax Data Flow Pipeline comes complete with all necessary storage, network, operating system and application software including Cloudera Data Flow (CDF). CDF is built on top of Apache Nifi, a powerful and user-friendly data routing application originally developed by the National Security Agency (NSA). With CDF, users have the ability to create simple graphical models that direct data across multiple networks with all necessary data provenance required for the most secure environments.  

A Truly Turn-key Platform 

According to Alex Lesser, PSSC Labs Vice President, the real value of CyberRax Data Flow Pipeline is a production-ready platform out of the box. “What we see happening with the delivery of CyberRax Data Flow Pipeline is a significant reduction in time and cost, allowing federal agencies to see true value much sooner,” says Lesser. With 48 processor cores, 4 GB memory per core, 264TB RAW storage capacity and dual 10GigE network backplane, CyberRax Data Flow Pipeline can fulfill the most demanding environments. Federal agencies can leverage the engineering expertise of PSSC Labs high performance and high reliability platform. That’s why it only makes sense that government agencies including NASA, NIH, DOD and many more have partnered with PSSC Labs in an effort to quickly process, store, analyze and protect their massive amounts of data. 

The U.S. Air Force recently deployed several CyberRax Data Flow Pipelines for both their NIPR and SIPR networks. The goal of these deployments is to collect, transport, and distribute cybersecurity data from every Air Force installation worldwide on both NIPR and SIPR. This solution allows for data to be sent to multiple core locations and fed downstream to multiple cybersecurity tools for analysis. 

American Made

As an American manufacturer, PSSC Labs provides a level of security and trust that other foreign manufacturers simply lack. The CyberRax Data Flow Pipeline can be purchased via several contract vehicles including GSA, CHESS and NETCENTS. PSSC Labs is already working closely with Department of Defense agencies to help them achieve their operational mandate to ingest and direct data from disparate sources all with the required security provisions out of the box. We know that government servers need to be 100% operational at all time, without sacrificing speed, analytics, or security, which is why our team of engineers is available for support for the lifetime of your system.

About PSSC Labs 

For technology powered visionaries with a passion for challenging the status quo, PSSC Labs is the answer for hand-crafted HPC and Big Data computing solutions that deliver relentless performance with the absolute lowest total cost of ownership.  

 We are true innovators offering high performance computing solutions to solve the world’s most demanding problems. For 25+ years, organizations of all sizes and from a variety of sectors rely on PSSC Labs’ computing systems. We are proud to support many departments within the United States government, Fortune 500 companies, as well as small and medium-sized businesses.  

All products are designed and built at the company’s headquarters in Lake Forest, California. 

PSSC Labs to Exhibit HPC Clusters for Weather Modeling and Analyzing Wildfire Fuel Loads at American Meteorological Society Annual Meeting

By marketing@site-a.com

High Performance Computing manufacturer teams with Atmospheric Data Solutions to provide turn-key supercomputer solutions for the AMS community

LAKE FOREST, CALIF., JANUARY 06, 2018 /PRWEB/ — PSSC Labs, a developer of custom High Performance Computing (HPC) and big data computing solutions, today announced it will showcase a series of High Performance Computing (HPC) solutions that are custom-built for weather modeling and wildfire analysis at the American Meteorological Society Annual Meeting (AMS) January 7-11, 2018 at the Austin Convention Center in Austin, Texas in Booth #736 in the Main Hall.

This year, the 2018 AMS Annual Meeting theme is “Transforming Communication in the Weather, Water, and Climate Enterprise Focusing on Challenges Facing Our Sciences.” The theme aims to enhance the scientific conference with a focus on communication science and practice to ensure the strengthened success of the enterprise in the future.

Weather research is rapidly advancing aided by the technological evolution of High Performance Computing servers. PSSC Labs delivers hand-crafted HPC and big data computing solutions based on its PowerWulf HPC cluster that are customized for the AMS community. PowerWulf HPC Clusters include pre-configured and fully validated blocks with the latest Intel HPC technology and all the necessary hardware, network settings, and cluster management software prior to shipping. Such solutions are paramount for dependable information in forecasting and air quality assessments for environmental analysts. Through the use of powerful, turn-key HPC cluster solutions for weather modeling solutions, climate scientists are able to combat wild fires such as the ones seen recently in Southern California.

PSSC Labs offers a complete hardware/software solution through its partnership with Atmospheric Data Solutions, LLC (ADS) to deliver powerful and customized supercomputing solutions for ADS’ weather modeling products. The weather modeling solutions that ADS creates include high impact weather forecast guidance products, tailored regional wildfire forecast guidance products, and utility load and outage forecasts – all requiring analysis of a large quantity of data that demands high performance computing to maximize accuracy and maximize the number of times models can be run daily.

PSSC Labs PowerWulf HPC Cluster offers a reliable, flexible, high performance computing platform for a variety of applications in the following verticals besides meteorological sciences including: Design & Engineering, Life Sciences, Physical Science, Financial Services and Machine/Deep Learning.

Source: http://www.prweb.com/releases/2018/01/prweb15058987.htm

Locating the Serial Number on a PSSC Labs Server Remotely

By marketing@site-a.com

The PSSC Labs support team is here to help you through troubleshooting your hardware and software related issues, but we will need to know a little bit about your computer in order to be truly effective.  To help us with this, please locate your PSSC Labs serial # and have it available when contacting the support team.  The serial # will quickly provide us with the hardware and software configuration, as well as any relevant warranty information.  The PSSC Labs Serial # is 6-digits and will start with a 3.  For example: 379884

The first method to use when looking for the Serial # is to see if it is hard coded into the machines DMI.  To determine this, run “dmidecode -t 1 | grep Serial”, which should output a Serial #.  If the Serial # is not programmed into the DMI, you will need to locate the Serial # physically on the machine chassis itself.

Depending on the type of PSSC Labs computer you have, the Serial # may be in a few different locations.  For most 1U rackmounted chassis we provide, the Serial # will be placed on a screw-on tab on the front-left of the machine.  Simply unscrew the thumbscrew and pull the tab out to reveal a label with the Serial # on it.  For most 2U and larger rackmounted machines, as well as desktop PowerStations, the label with the Serial # is usually easily located on the rear of the machine.

If you are having difficulty locating the Serial #, please let us know first thing when contacting support so we can assist you with locating it.  

PSSC Labs Launches PowerWulf HPC Clusters With Pre-Configured Intel Data Center Blocks

By marketing@site-a.com

Built with Intel Scalable CPUs and Omni-Path to address cloud, HPC and business critical workloads

LAKE FOREST, Calif., Dec. 11, 2017 /PRNewswire/ — PSSC Labs, a developer of custom HPC and Big Data computing solutions, today announced its PowerWulf HPC clusters are now available with Intel’s new Xeon® Scalable Processors and Intel’s Omni-Path HPC Fabric to deliver the performance needed to tackle cutting edge computing tasks including real-time analytics, virtualized infrastructure and high-performance computing. 

PowerWulf clusters are built with Intel’s Data Center Blocks to ensure a truly turnkey solution that addresses customer integration challenges. Today’s customer datacenters require unique server solutions that run complex, business-critical workloads. Intel Data Center Blocks configurations are purpose-built with all-Intel technology, optimized to address the needs of specific market segments. These fully validated blocks deliver performance, reliability and quality for solutions customer want and can trust to handle their demanding cloud, HPC, and business critical workloads.

PSSC Labs PowerWulf HPC Clusters are available as config-to-order (CTO) to meet the specific needs of a customer. Key features of these solutions include:

  • Pre-configured and fully validated blocks with the latest Intel HPC technology
  • Powered by the Intel Xeon processor Scalable family, delivers an overall performance increase up to 1.65x compared to the previous generation, and up to 5x Online Transaction Processing warehouse workloads versus the current install base.
  • 2 operating system options to choose: RedHat, SUSE, and CentOS Linux
  • Multiple models with different support options
  • Intel Fabric Suite 10.5.1, Lustre 2.10
  • Intel Omni-Path Host Fabric Interface (Intel OP HFI) Adapter 100 Series and FDR/EDR InfiniBand Fabric
  • Intel Datacenter SATA and NVMe Solid State Drives (SSD)

“Intel’s integrated and fully-validated Data Center Blocks enables PSSC Labs to deliver more efficient and turnkey approach and reduce time to market, complexity and the costs of system design, validation and integration,” said Alex Lesser, EVP of PSSC Labs. “Partnering with Intel allows us to offer our customers the latest hardware options in our line of custom turn-key PowerWulf HPC clusters for a variety of applications across government, academic and commercial environments.”

PowerWulf HPC clusters also feature PSSC Labs CBeST Cluster Management Toolkit (Complete Beowulf Software Toolkit) to deliver a preconfigured solution with all the necessary hardware, network settings and cluster management software prior to shipping. With its component structure, CBeST is the most flexible cluster management software package available.

Every PowerWulf HPC Cluster includes a three-year unlimited phone / email support package (additional year support available) with all support provided by PSSC Labs US-based team of engineers. PSSC Labs is an Intel HPC Data Center Specialist and has been a Platinum Provider with Intel since 2009. For more information see https://www.site-a.com/solutions/hpc-cluster/

Source: https://www.prnewswire.com/news-releases/pssc-labs-launches-powerwulf-hpc-clusters-with-pre-configured-intel-data-center-blocks-300569232.html