Scaling Machine Learning Systems for the Data Scientist

By marketing@site-a.com

The use of machine learning (ML) is on the rise. From a data scientist’s perspective, the computing challenge is how to scale the ingesting of more data in faster times to train machine learning algorithms, as well as how to scale processing power. Looking deeper, the issue is that ML applications parse ever-growing amounts of data requiring enormous parallel processing capabilities using large numbers of cores.

Traditional computing systems based on standard CPUs will not suffice. They cannot process the data, train the machine learning algorithms, or run the ML applications against new data in an efficient manner. In most cases, scaling legacy systems to the processing level required is too costly. Even if the investment is made, the time it takes to train and run ML applications is impractical for the needs of the business.

What’s needed is an infrastructure update that delivers the required parallel processing performance at a reasonable cost. In most cases, the best solution is one that combines multithreaded CPUs with GPUs, large memory, high-performance interconnects, and HPC storage solutions with high I/O and throughput, plus low latency features.

What’s Different?

Organizations have been scaling their systems for things like Big Data and the use of more sophisticated analytics for years. In most cases, solutions based on traditional CPU architectures were enough.

Why is machine learning different? Why do these traditional solutions not fit the bill?

Many organizations found their installed systems were hitting a wall because of the amounts of data involved and the nature of the ML algorithms. Training models took too long to run due to computing limitations.

When confronted with this problem, organizations have looked for systems that lend themselves to the requirements of ML and machine learning algorithms. Such systems often include HPC servers with greater processor performance, systems that scale up (vs. scale-out), I/O solutions with bandwidth, and accelerator technologies such as GPUs or FPGAs.

Solutions with these characteristics get to the heart of the problem for ML / Machine Learning — core starvation. CPUs are designed for serial processing. Machine learning training and applications must be done in parallel on many more cores than CPUs can provide. Accelerators overcome this problem. GPUs offer thousands of cores, and custom-designed processors (ASIC, FPGAs) complement CPU processing capabilities.

Such accelerators offer a massively parallel architecture that economically delivers the needed parallel compute performance. Going hand-in-hand with the use of accelerators, systems that can scale to meet the demands of ML applications also must have high-speed interconnects, increased memory size, and fast storage.

Technology from an Experienced Partner

PSSC Labs has delivered tens of thousands of custom-engineered HPC servers to higher education, government agencies, small/medium businesses, and large enterprise organizations across 36 countries.

PSSC Labs delivers integrated HPC / High Performance Computing solutions for ML that tightly integrate and optimize hardware and software. Its PowerServe Uniti Servers include the latest components from Intel® and Nvidia®.

PSSC Labs GPU options include:

  • NVIDIA Tesla P100 GPU accelerators for PCIe based servers. Tesla P100 with NVIDIA NVLinkdelivers up to a 50X performance boost for the top HPC applications and all deep learning frameworks.
  • NVIDIA Tesla V100 Tensor Core, powered by NVIDIA Volta architecture, is a data center GPU to accelerate HPC and AI workloads.
  • NVIDIA T4 GPU, which accelerates cloud workloads, is used for HPC, deep learning training, and inference, machine learning, and data analytics.
  • GEFORCE RTX 2080 Ti is NVIDIA’s flagship graphics card based on NVIDIA Turing™ GPU architecture and ultra-fast GDDR6 memory.

Systems that use these and other GPUs to scale ML workloads need high-performance interconnect technologies to make cost-effective use of their performance capabilities. Interconnect technologies available include InfiniBand, Omni-Path, and remote direct memory access (RDMA).

Most important, PSSC Labs’ family of HPC systems are Machine Learning ready systems that are purpose-built for an organization’s needs. The solutions come production ready, which is known to be critical from a data scientist’s perspective. PSSC Labs does not need to spend time on IT issues, as clients are assured the systems they are using for their ML efforts can scale to meet the demands of their applications. 

To learn more about our HPC systems specifically designed to meet the need of data scientists across various industries, click the button below to schedule a meeting with one of our knowledgable Solutions Architects. 


Schedule a Meeting

Your Parallel File System is Evolving

By marketing@site-a.com

Distributed parallel file systems have been a core technology to accelerate high-performance computing (HPC) workloads for nearly two decades (Lustre 1.0.0 was released in 2003). While first the purview of supercomputing centers, distributed parallel file systems are now routinely used in mainstream HPC applications.

However, if you have not looked at the available options in a while, you will find new thinking is required when evaluating solutions for today’s workloads. What’s changed? Almost everything.   

First, today’s workloads and the data they operate on are different. Applications for modeling physical systems, Big Data analytics, and the training and use of artificial intelligence (AI) and machine learning (ML) typically have very low latency and high IO throughput requirements. They need a storage and interconnectivity infrastructure that efficiently stages data for processing and blasts that data to servers for processing without delays. Solutions must have both very high sustained throughputs and low latency.

Another workload and data change that impacts the choice of a file system is that many applications now must accommodate both large blocks of data and many small files. For instance, training an ML model might use millions of small files, while a Big Data analytics workload runs on one massive dataset. There also is much more use of metadata (consisting of numerous small files) in many workloads today.

Second, the processors used on today’s workloads are different. In the past, cluster architecture came down to a choice between fat nodes (typified by their large number of cores per node) or thin (sometimes called skinny) nodes with fewer, faster-performing cores per node. The choice of one type of node versus another, based on the nature of the workload, had great implications for a file system choice.

Advanced modeling and simulation algorithms, such as a finite element simulation, are massively parallel workloads. They often required fat nodes because a job needs to use a very large number of cores at a given time. Such workloads where job execution steps are run in parallel typically need a fast and low latency network and storage with high IO and throughput.

Sequential workloads were each step of a job can be executed independently of the others benefit from very fast CPUs. Skinny nodes are well-suited for this type of workload. From an infrastructure perspective, storage systems for a cluster built of skinny nodes must have the capacity to handle and manage the typically large datasets analyzed and processed in workloads that run sequentially. Additionally, the storage systems must be able to support the numerous read/write operations of many small-sized files without degradation in performance.

In both cases, the file system selected had to support the performance requirements of the applications and data types. Now, there is a relatively new processor choice that places different demands on a file system. That new option is the growing use of GPUs for compute-intensive applications (versus their traditional use for rendering). GPUs include many cores that execute in parallel. They are typically expensive. Making efficient use of a GPU’s capacity requires a file system that can ensure the cores are satiated.

Finally, the last difference is that the available distributed parallel file systems have changed. Some have been around for many years but now have new names. All are frequently updated with new features and capabilities. Today, the three main choices are:

Lustre: Lustre file system software is available under the GNU General Public License (version 2 only) and provides high-performance file systems for computer clusters ranging in size from small workgroup clusters to large-scale, multi-site clusters. It’s most recent update is Lustre 2.13, which was released on December 5, 2019. This release added a new performance-related features Persistent Client Cache (PCC), which allows direct use of NVMe and NVRAM storage on the client nodes while keeping the files part of the global filesystem namespace, and OST Overstriping, which allows files to store multiple stripes on a single OST to better utilize fast OSS hardware. Also, in this release, the PFL functionality was enhanced with Self-Extending Layouts (SEL) to allow file components to be dynamically sized, to better deal with flash OSTs that may be much smaller than disk OSTs within the same filesystem.

IBM Spectrum Scale: Formerly known as IBM General Parallel File System (GPFS), IBM Spectrum Scale is a high-performance clustered file system software developed by IBM. It can be deployed in shared-disk or shared-nothing distributed parallel modes. Features of its architecture include distributed metadata (including the directory tree), efficient indexing of directory entries for very large directories, and distributed locking. Additionally, filesystem maintenance can be performed while systems are online, ensuring the filesystem is available more often, thus keeping the HPC cluster itself available longer.

BeeGFS: BeeGFS is a parallel file system that was initially developed at Fraunhofer Center for High Performance Computing in Germany by a team led by Sven Breuner, who later became the CEO of ThinkParQ, the spin-off company that was founded in 2014 to maintain BeeGFS. BeeGFS combines multiple storage servers to provide a highly scalable shared network file system with striped file contents. This way, it allows users to overcome the tight performance limitations of single servers, single network interconnects, and a limited number of hard drives. In such a system, high throughput demands of large numbers of clients can easily be satisfied, but even a single client can benefit from the aggregated performance of all the storage servers in the system.

Gluster: Similar to the way the Lustre name was derived combining Linux and cluster, Gluster combines GNU and cluster. GlusterFS is a scalable network filesystem suitable for data-intensive tasks such as cloud storage and media streaming. The GlusterFS architecture aggregates compute, storage, and I/O resources into a global namespace. Capacity is scaled by adding additional nodes or adding additional storage to each node. Performance is increased by deploying storage among more nodes. High availability is achieved by replicating data n-way between nodes.

Making Sense of the Changes and Choices

Given these changes in workloads, processor choices, and available parallel file systems, how do you match a system to your company’s HPC needs? You can certainly do it yourself. However, with HPC analytics and AI applications becoming mainstream, most companies do not have the time or in-house expertise to evaluate, select, assemble, deploy, and maintain systems from scratch.

A better approach is to team with an industry partner that has expertise in the new applications and solutions, plus best practices developed from a long track record of successful deployments. That’s where we can help.

PSSC Labs has a more than 30 years history of delivering systems that meet the most demanding workloads across industries, government, and academia.

Its offerings include the Parallux Storage Clusters, which are a cost-efficient and scalable storage solution. Key features include:

  • Scale to tens of petabytes of capacity with no downtime upgrade
  • Storage tiers available for warm/cold data requirements
  • Distributed metadata for high performance and redundancy
  • Compatible with POSIX File Systems
  • Extreme performance exceeding 10GB/sec sustained IO
  • Factory installation of all necessary hardware, software, and networking components
  • Only enterprise-grade components used for maximum reliability and performance

This solution and other PSSC Labs systems are designed to meet the compute requirements of modern enterprise applications today. Such systems will increasingly become more important as companies make greater use of analytics on more and more datasets, as well as embracing AI and ML.

Marrying Big Data Analytics and Supercomputing

By marketing@site-a.com

The increased reliance by mainstream companies on analytics and artificial intelligence (AI) for business intelligence and business process management chores is driving a need for a type of system that integrates Big Data analytics and supercomputing capabilities.

Assembling a system for one or the other (Big Data analytics or HPC and supercomputing) is challenging enough; getting all the needed elements into one optimized system is even harder. To put the issue into perspective, consider the characteristics of systems optimized for one versus the other.

Key features of HPC and supercomputing systems include:

  • Scalable computing with high bandwidth, low-latency, global memory architectures
  • Tightly integrated processor, memory, interconnect, and network storage
  • Minimal data movement (accomplished by loading data into memory)

In contrast, key characteristics of a Big Data analytics system typically include:

  • Distributed computing
  • Service-Orientated Architecture
  • Lots of data movement (processes include sorting or streaming all the data all the time
  • Low-cost processor, memory, interconnect, and local storage

Factors Impacting Architectural Choices

One of the biggest challenges is matching compute, storage, memory, and interconnection capabilities to workloads. Unfortunately, the only constant is that AI and Big Data analytics workloads are highly variable. Some factors to consider include:

AI Training

Companies need enormous compute and storage capacity to train AI models. The demand for such capabilities is exploding. Since 2012, the amount of compute used in the largest AI training runs has been increasing exponentially with a 3.4-month doubling time.

Given that with Moore’s Law compute power doubles in 18 months, how is it possible to deliver such compute capacity and keep pace as requirements grow? The answer, as most likely know, is to accelerate these workloads with field-programmable gate arrays (FPGAs) and graphics processing units (GPUs) that complement a system’s CPUs.  Intel® and Nvidia® are both staking their claim with new FPGA and GPU products that are becoming standards. Nvidia® Tesla and RTX series GPUs are synonymous with AI and machine learning.  Intel®’s recent acquisition of Altera has jump started their entry into this market. 

Where do these elements play a role? Deep learning powers many AI scenarios, including autonomous cars, cancer diagnosis, computer vision, speech recognition, and many other intelligent use cases. GPUs accelerate deep learning algorithms that are used in AI and cognitive computing applications.

GPUs help by offering thousands of cores capable of performing millions of mathematical operations in parallel. Like a GPU’s use for graphic rendering, GPUs for deep learning deliver a significant number of matrix multiplication operations per second. 

Big Data Ingestion and Analysis

System requirements can vary greatly depending on the type of data being analyzed. Running SQL queries against a structured database has vastly different systems requirement than say real-time analytics or cognitive computing algorithms on streaming data from smart sensors, social media threads, or clickstreams.

Depending on the data and the analytics routines, a solution might include a data warehouse, NoSQL database, graphic analysis capabilities, or Hadoop or MapReduce processing. (Or a combination of several of these elements.)

Beyond architecting a system to meet the demands of the database and analytic processing algorithms, attention must be paid to data movement. Issues to be considered include how to ingest the data from its original source, stage data in preparation for analysis, and ensure processes are satiated to make efficient use of the expensive cores.

Optimized, Integrated Solutions

Vendors and the HPC community are working to address the Big Data challenge in a variety of ways – especially with the general acceptance of AI and its dependence on large data sets.

A suitable system must be able to be dynamically provisioned by the users to handle different data workflows, including databases (both relational database systems and NoSQL style databases), Hadoop/HDFS based workflows (including MapReduce and Spark), and more custom workflows perhaps leveraging a flash-based parallel file system.

The challenge for most companies is that they might have system expertise in HPC or Big Data, but often not in both. As such, there may be a skills gap that makes it hard to bring together the essential elements and fine-tune the system’s performance for critical workloads.

To overcome such issues, many companies are looking for turnkey solutions that combine the needed processing, storage, memory, and interconnect technologies for Big Data analytics and HPC/supercomputing into a tightly integrated system.

Delivering such a solution requires expertise and real-world best practices across both HPC and Big Data domains, plus deep industry knowledge about the specific data sources and analytics applications.

PSSC Labs has a more than 30 years history of delivering systems that meet the most demanding workloads across industries, government, and academia.

Its offerings include the PowerServe Uniti Server and PowerWulf ZXR1+ HPC Cluster lines that leverages the latest components from Intel® and Nvidia®. These servers are ideal for applications including AI and deep learning; design and engineering; life sciences including genomics and drug discovery and development; as well as for computational and data analysis for chemical and physical sciences.

PSSC Labs also offers CloudOOP Big Data Servers and CloudOOP Rax Big Data Clusters that deliver the highest level of performance one would expect in an enterprise server combined with the cost-effectiveness of direct attach storage for Big Data applications. The servers are certified compatible with Cloudera, MapR, Hortonworks, Cassandra, and Hbase. And deliver 200+ MB/sec sustained IO speeds per hard drive (which is 30%+ faster than other OEMs.)

These solutions and other PSSC Labs systems are designed to meet the requirements of modern Big Data analytics in enterprises today. Such systems will increasingly become more important as companies explore new ways to make use of the explosive amounts of data to improve operations, better engage customers, and quickly react to new business opportunities.

The bottom line is that the world is entering an era that requires extreme-scale HPC coupled with data-intensive analytics. Companies that succeed will need systems that bring essential compute, storage, memory, and interconnect elements together in a manner where each element complements the others to deliver peak performance at a nominal price.

The Role of Low Latency File Access in Accelerating AI Workloads

By marketing@site-a.com

The use of artificial intelligence (AI) is rapidly moving from the lab into the mainstream. The reason? Businesses believe AI can deliver operational cost savings, improve decision making, enhance customer interactions, speed data mining, and boost data security. As such, the number of companies using AI has grown by 270% in the past four years. As a result, organizations need to design high-performance computing architectures for AI workloads. 

Supporting AI efforts requires high-performance computing (HPC) capabilities to perform rapid analysis, tune neural net models, and conduct machine learning by examining large datasets. Fortunately, HPC requirements for AI are similar to other compute-intensive applications (e.g., Big Data analytics, forecasting, modeling, and finite element simulations) that are also increasingly being introduced into the enterprise today. That means there are many high-performance core compute, storage, and networking technologies available, which have made their way from supercomputing centers and academic labs into the enterprise.

However, several factors determine what type of infrastructure elements are needed for specific AI applications. Many AI efforts need speedy execution. That is the case for AI applications that do things like power autonomous systems, engage customers in real-time via chat or natural language, or identify outliners to prevent fraud. Such applications must run their analyses and get actionable information in real-time.

Selecting the right technology solution

To achieve the necessary performance, AI deployments typically make use of expensive GPU processing arrays. For workloads to run cost-effectively, there is a need to support high data rates to keep the processors satiated. That, in turn, dictates the use of ultrafast interconnect technology and tightly coupled high-performance storage.

Looking deeper, some of the GPU options include:

  • NVIDIA Tesla P100 GPU accelerators for PCIe based servers. Tesla P100 with NVIDIA NVLink delivers up to a 50X performance boost for the top HPC applications and all deep learning frameworks.
  • NVIDIA Tesla V100 Tensor Core, powered by NVIDIA Volta architecture, is a data center GPU to accelerate HPC and AI workloads.
  • NVIDIA T4 GPU, which accelerates cloud workloads, is used for HPC, deep learning training, and inference, machine learning, and data analytics.
  • GEFORCE RTX 2080 Ti is NVIDIA’s flagship graphics card based on NVIDIA Turing™ GPU architecture and ultra-fast GDDR6 memory.

Systems that use these and other GPUs to accelerate AI workloads need high-performance interconnect technologies to make cost-effective use of their performance capabilities. The internet technologies of choice include InfiniBand, Omni-Path, and remote direct memory access (RDMA). 

What are their capabilities?

  • InfiniBand is a computer-networking communications standard used in HPC systems that features very high throughput and very low latency. It is used as either a direct or switched interconnect between servers and storage systems, as well as an interconnect between storage systems.
  • Omni-Path(also Omni-Path Architecture or OPA) is a high-performance communication architecture from Intel. It delivers low communication latency and high throughput.
  • RDMA is an industry-standard that supports what is known as zero-copy networking by enabling the network adapter to move data directly to or from the application. This eliminates both the operating system and CPU involvement, so it is exceptionally faster than other solutions.

The plethora of GPU and interconnect technologies choices is a double-edged sword. On the plus side, the right combination will produce an optimized system to accelerate a specific AI application. On the downside, many businesses do not have expertise in these technologies and need help selecting the best solution for their application and optimizing a system’s performance.

When selecting elements in a system, it’s critical to determine which interconnect solution provides low latency file access to help AI and HPC workloads achieve higher performance and scalability.

These challenges exist when trying to configure a system for any HPC application, but the issues are especially important with AI applications. They need fast access to data to reduce training time in a deep learning scenario but also in supporting fast decision making in production environments.

Determining which is best for you

So how do you determine which storage is best for an AI application? Beyond basics like determining cost/performance issues when using hard drives versus solid-state and flash drives, there are storage file system and architecture issues to consider. Do you use a distributed architecture? Do you need a parallel file system? The bottom-line: AI applications need storage solutions that offer the highest throughput, lowest latency data access for CPU- and GPU-intensive AI, and HPC workloads.

In the final analysis, to optimize the running of AI workloads and make the most efficient use of expensive GPU arrays, compute solutions must bring together the right GPU, high-performance storage, and interconnect technologies. These technologies must be tightly integrated and tuned to optimize the solution’s performance when running AI workloads.

Modern Threat Hunting: Uncovering Hidden Indicators of Compromise

By marketing@site-a.com

Modern threats evolve rapidly and are increasing in frequency and malevolence. Detection delays can represent millions in damage, jeopardize corporate brands, and put security jobs in jeopardy. Are you well positioned to adapt to these new-age threats? Relying on threat data that is stale or inaccurate is a recipe for failure. 

This webinar will bring you up to date on how threats are changing and new, proven methods for identifying and countering them early. Workflows and application use will be used to present critical path methods for threat hunting and preemptive breach discovery.

Watch and learn

  • The changing nature of threats and organizations’ growing exposure
  • How data science and cognitive approaches are driving more preemptive discovery of modern attacks
  • The key role Open Source will play in ensuring you have the latest protections against the newest threats
  • Why automation and visualization are absolutely necessary to deal with exploding attack volume and sophistication

Reduce your risk by knowing how threats are evolving and the best methods for countering them.