Infrastructure Considerations for Containers and Kubernetes

By marketing@site-a.com

Containers and Kubernetes are at the heart of a broad industry shift where applications and services are based on a microservices architecture. Specifically, microservices are being rapidly adopted as a means of building and modernizing distributed applications allowing them to be more scalable, flexible, resilient, and easier to build.

Instead of building self-contained, monolithic applications, a microservices approach breaks applications into modular, independent components that can be dynamically integrated with one another using application programming interfaces (APIs).

Increasingly, companies are using containers to power their microservices application architectures. Containers encapsulate a lightweight runtime environment for an application. Specifically, containers enable finer-grained execution environments, permit application isolation, and are lightweight.

Furthermore, containers include everything needed to run, such as code, dependencies, libraries, binaries, and other elements. Today, Docker is the most popular choice for building and running containers.

Compared to Virtual Machines, containers share the OS kernel instead of having a full copy of it and take up less space. Because they do not require OS spin-up time associated with a VM, containers initialize faster. In general, containers start in seconds or even milliseconds, which is much faster than VMs. As such, containers deliver performance characteristics that match the needs of a microservices architecture. In particular, the quick instantiation maps better to the unpredictable workload characteristics associated with microservices. 

The growing embracement of containers was validated in a 2019 industry container usage survey that found the median number of containers per host doubled (to 30) between 2018 and 2019. And the maximum per-node density was 250 containers, which was a 38% increase from 2018.

Managing Your Containers

With such explosive growth in the use of containers, companies need a way to oversee and manage their efforts. That’s where Kubernetes comes in.

Kubernetes is an open-source container orchestrator system for automating deployment, scaling, and management of application containers across clusters of hosts. It was originally designed by Google and is now maintained by the Cloud Native Computing Foundation. Kubernetes works with a range of container tools, including Docker. It groups containers that make up an application into logical units for easy management and discovery.

Kubernetes provides a framework to run distributed systems resiliently. It takes care of scaling and failover for an application. For example, in a production environment, Kubernetes can start a new container if one goes down. Thus, helping to ensure there is no application downtime. Additionally, Kubernetes provides service discovery and load balancing, storage orchestration, automated rollout and rollbacks, self-healing features, configuration management, and more.

While there are a handful of container orchestrators available today, Kubernetes dominates the market. In addition to the widely used open-source variant, some commercial offerings such as Red Hat OpenShift are built on Kubernetes. (The commercial offers add enterprise features and support.)

Kubernetes can be deployed on a bare-metal cluster or on a cluster of virtual machines. Kubernetes, in turn, can orchestrate the containers it manages directly on bare metal or on virtual machines. Most instances of Kubernetes today are run on VMs running on-premises or in the cloud.

Bare-metal instances are not as common. However, there are use cases where they offer advantages. For example, a network edge application might be too latency-sensitive to tolerate the overhead created by a VM. Or an application (such as machine learning) might need to run on GPUs or other hardware accelerators, which do not lend themselves to VMs.

Optimized, Integrated Solutions

Running container workloads comes down to hardware. Businesses need physical machines, with CPUs, memory, and local persistent storage. In addition, they need some shared persistent storage and networking element to hook up all the machines.

A suitable system must be able to be dynamically provisioned by the users to handle different data workflows. Many companies are looking for turnkey solutions that combine the needed processing, storage, memory, and interconnect technologies to provide either the bare metal or VM foundation for their container and microservices efforts. Delivering such a solution requires expertise and real-world best practices across both HPC and container/Kubernetes domains, plus deep industry knowledge about the specific applications.

PSSC Labs has a more than 30 years history of delivering systems that meet the most demanding workloads across industries, government, and academia. Its offerings include the PowerServe Uniti Servers line that leverages the latest components from Intel® and Nvidia®. These servers are ideal for a wide range of applications, including AI and deep learning, as well as for computational and data analysis.

PSSC Labs also offers CloudOOP Big Data Servers that deliver the highest level of performance one would expect in an enterprise server combined with the cost-effectiveness of direct attach storage for Big Data applications. The servers deliver 200+ MB/sec sustained IO speeds per hard drive (which is 30%+ faster than other OEMs.)

As the number of containers per host grows and Kubernetes use grows, these solutions and other PSSC Labs systems are designed to meet the requirements of enterprises today. Such systems will increasingly become more important as companies explore new ways to make use of the containers, Kubernetes, and microservices to serve their users better and quickly react to new business opportunities.

Supercomputing for Design & Engineering SMBs

By marketing@site-a.com

Supercomputing is most commonly thought of as a tool only for the largest government agencies, top commercial companies, and most prestigious universities — not your everyday small and medium-sized businesses (SMBs). The focus of supercomputing is typically on the largest systems, which scale well beyond many thousands of processor cores. Many SMBs believe supercomputing capabilities to be well beyond their reach, but nothing could be further form the truth. SMBs need a competitive advantage to obtain government and commercial contracts, and that advantage is a powerful, scalable, high performance supercomputing platform.

Over the past 12 months, interest and deployment of high performance computing clusters at SMBs performing computational fluid dynamics (CFD), finite element analysis (FEA) and structural engineering work have trended significantly upward. Without HPC systems, the work these organizations are accomplishing would be nearly impossible. But with the right systems, it could be the lifeline to their organizational success. “Supercomputing is not just for the large and powerful organizations,” says Alex Lesser, Vice President of PSSC Labs. “Supercomputing is an accessible tool for everyone. Small and Medium Businesses need to leverage these tools to ensure their success, and it does not have to be costly or intimidating.”

PSSC Labs works closely with several trusted leaders in the design and engineering space. These organizations develop state-of-the-art physics-based models for test programs, modeling and simulation analysis solutions for highly dynamic events and reactivity of energetic materials, and much more. Many of our clients also design simulation software in multiple areas of physics, including computational fluid dynamics (CFD), aero and hydro acoustics, and more. The technology these best-of-breed design and engineering firms have developed are applicable to many areas of interests to scientists, engineers, technologists, and educators. We’re proud to support game-changing organizations with the necessary HPC hardware systems they need to build the best of the best, while giving them immediate, frontend access to a variety of CFD applications, including Ansys Fluent, Star CCM, CFD++, and OpenFOAM, as well as Finite Element Analysis (FAE) applications, like Abaqus FEA, COMSOL Multiphysics and SimScale.

POWERWULF ZXR1+ HPC CLUSTER

Starter configuration for under $100k

  • 200 Intel or AMD Processor Cores
  • 4 GB of Memory Per Processor Core
  • 100 Gbps High Speed Network
    (Intel Omnipath or Mellanox Infiniband)
  • 10 TB High Performance Flash Storage
  • 40 TB Long Term / Secondary Storage
  • CBeST Cluster Management Toolkit
    • Linux Cluster Operating System
    • Message Passing Libraries
    • Batch Scheduled
  • Complete 3 Year Service Level Agreement

engineering hpc cluster

While some SMBs might be hesitant to deploy an on-premise HPC platform due to the lack of an internal IT department or perceived level of complexity, these concerns are easily put to rest when partnering with the right partner. PSSC Labs’ goal is to deliver the most turn-key, highest performance, head-ache free HPC experience possible. Our PowerWulf ZXR1+ Clusters include all necessary hardware, software, and networking integrated by 20+ year HPC experts. “We are very proud of our 57 step testing and integration process, which is considered one of the industry’s most rigorous,” begins Larry Lesser, PSSC Labs Chief Technology Officer. “The results of this painstaking process are evident with a record of never delivering a DOA system and a consistently proven 99.99% system uptime.”  PSSC Labs stands firmly behind their product with a complete hardware and software service level agreement for up to five years. Because PSSC Labs is both the manufacturer of the hardware and developer of system software, SMBs have one phone number to call for immediate access to knowledgeable support.

Cloud providers and brokers will continue to try scaring SMBs into believing that they are incapable of supporting an on-premise HPC solution. They may argue the merits of using offerings from Amazon, Microsoft and Google will save companies significant cost, offer superior performance and greater scalability, but they’re wrong. SMBs in the Design & Engineering space are comprised of very intelligent and experienced engineers who can evaluate the numbers on their own and are not easily swayed. Companies who choose the cloud over on-premise models are accepting ever-expanding monthly bills, loss of control and governance, and lower security, while allowing their data to be completely at the mercy of a third party. With the threat of cyber attacks ever-looming, putting critical and confidential data on the cloud is possibly one of the most irresponsible things SMBs could do. A quick read of the “Cloud Hopper” investigation by the Wall Street Journal should serve as a stark warning.

More and more SMBs are deploying HPC systems than ever before. On-premise, high-performance supercomputers are well within the grasp of design & engineering organizations of all sizes. SMBs need to utilize these tools to develop new products, win more contracts and push their businesses forward.

The Role of Low Latency File Access in Accelerating AI Workloads

By marketing@site-a.com

The use of artificial intelligence (AI) is rapidly moving from the lab into the mainstream. The reason? Businesses believe AI can deliver operational cost savings, improve decision making, enhance customer interactions, speed data mining, and boost data security. As such, the number of companies using AI has grown by 270% in the past four years. As a result, organizations need to design high-performance computing architectures for AI workloads. 

Supporting AI efforts requires high-performance computing (HPC) capabilities to perform rapid analysis, tune neural net models, and conduct machine learning by examining large datasets. Fortunately, HPC requirements for AI are similar to other compute-intensive applications (e.g., Big Data analytics, forecasting, modeling, and finite element simulations) that are also increasingly being introduced into the enterprise today. That means there are many high-performance core compute, storage, and networking technologies available, which have made their way from supercomputing centers and academic labs into the enterprise.

However, several factors determine what type of infrastructure elements are needed for specific AI applications. Many AI efforts need speedy execution. That is the case for AI applications that do things like power autonomous systems, engage customers in real-time via chat or natural language, or identify outliners to prevent fraud. Such applications must run their analyses and get actionable information in real-time.

Selecting the right technology solution

To achieve the necessary performance, AI deployments typically make use of expensive GPU processing arrays. For workloads to run cost-effectively, there is a need to support high data rates to keep the processors satiated. That, in turn, dictates the use of ultrafast interconnect technology and tightly coupled high-performance storage.

Looking deeper, some of the GPU options include:

  • NVIDIA Tesla P100 GPU accelerators for PCIe based servers. Tesla P100 with NVIDIA NVLink delivers up to a 50X performance boost for the top HPC applications and all deep learning frameworks.
  • NVIDIA Tesla V100 Tensor Core, powered by NVIDIA Volta architecture, is a data center GPU to accelerate HPC and AI workloads.
  • NVIDIA T4 GPU, which accelerates cloud workloads, is used for HPC, deep learning training, and inference, machine learning, and data analytics.
  • GEFORCE RTX 2080 Ti is NVIDIA’s flagship graphics card based on NVIDIA Turing™ GPU architecture and ultra-fast GDDR6 memory.

Systems that use these and other GPUs to accelerate AI workloads need high-performance interconnect technologies to make cost-effective use of their performance capabilities. The internet technologies of choice include InfiniBand, Omni-Path, and remote direct memory access (RDMA). 

What are their capabilities?

  • InfiniBand is a computer-networking communications standard used in HPC systems that features very high throughput and very low latency. It is used as either a direct or switched interconnect between servers and storage systems, as well as an interconnect between storage systems.
  • Omni-Path(also Omni-Path Architecture or OPA) is a high-performance communication architecture from Intel. It delivers low communication latency and high throughput.
  • RDMA is an industry-standard that supports what is known as zero-copy networking by enabling the network adapter to move data directly to or from the application. This eliminates both the operating system and CPU involvement, so it is exceptionally faster than other solutions.

The plethora of GPU and interconnect technologies choices is a double-edged sword. On the plus side, the right combination will produce an optimized system to accelerate a specific AI application. On the downside, many businesses do not have expertise in these technologies and need help selecting the best solution for their application and optimizing a system’s performance.

When selecting elements in a system, it’s critical to determine which interconnect solution provides low latency file access to help AI and HPC workloads achieve higher performance and scalability.

These challenges exist when trying to configure a system for any HPC application, but the issues are especially important with AI applications. They need fast access to data to reduce training time in a deep learning scenario but also in supporting fast decision making in production environments.

Determining which is best for you

So how do you determine which storage is best for an AI application? Beyond basics like determining cost/performance issues when using hard drives versus solid-state and flash drives, there are storage file system and architecture issues to consider. Do you use a distributed architecture? Do you need a parallel file system? The bottom-line: AI applications need storage solutions that offer the highest throughput, lowest latency data access for CPU- and GPU-intensive AI, and HPC workloads.

In the final analysis, to optimize the running of AI workloads and make the most efficient use of expensive GPU arrays, compute solutions must bring together the right GPU, high-performance storage, and interconnect technologies. These technologies must be tightly integrated and tuned to optimize the solution’s performance when running AI workloads.

Top Tips to Improving Your System Administration Skills

By marketing@site-a.com

You Envision the Future

LAKE FOREST, CALIF., SEPT 12, 2019 — PSSC Labs, a developer of custom High Performance Computing and Big Data computing solutions today shares their top tips for improving your system administration skills.

At PSSC Labs, our mission is to deliver reliable, high performance turnkey platforms that help organizations like yours perform necessary data tasks without giving up control of your data.  That means keeping your system administration skills up-to-date is more important than ever! Of course, we’re always here to help you should you ever encounter any issues. In fact, you can contact our Support line at any time and we will be happy to assist you. But just to give you a little boost, we’ve outlined the top five support issues that on-premise model users encounter, and what to do if it happens to you.

Machine Not Powering On

PSSC Labs produces turnkey platforms, so they’re ready to be plugged in and get going upon delivery. Should you ever notice that one of your machines isn’t powering on, the first thing you should check is power supply. Verify that the power supply is receiving power from the source by verifying that the power connectors are securly plugged in on both ends.

If this doesn’t work, try a new power port on the PDF and then move on to a new power cable. If your machine still does not power on you can try using a power supply from your spare parts kit (if a spare parts kit was provided with your order) or from another working machine.  Machines with faulty memory can also exhibit this type of behavior, so shuffling or re-seating the memory is also worth trying. 

Machine Has Crashed or Rebooted

On-premise models can crash or reboot for various reasons. They’re often accompanied by BMC System Event Log messages.  Start by checking the output of the command ‘ipmitool sel list’ and check “/var/log/messages” for errors correlating to the date / time of the crash.

Updating your machine’s BIOS to the latest version is recommended as well, as it can fix crashing systems as well as improve the machine’s overall performance.  Verify that the CPUs are not overheating by checking the “sensors” command while the machine is under load.  Faulty memory can also cause crashes, so re-seating or shuffling memory can help to resolve these issues as well.  The causes of system crashes are sometimes hard to pin down, so do not hesitate to contact the PSSC Labs support team for troubleshooting assistance.

Software Support

PSSC Labs solutions often come pre-installed with multiple software applications, ranging from job submission software to parallel computing platforms, from weather modeling to data analysis.

Many of these packages are third party applications and not created by PSSC Labs, so if you ever have an issue with pre-loaded software, it’s recommended to first contact the software developers. If you still need assistance afterwards, contact the PSSC Labs support staff who can promptly assist you further.

Machine is Reporting Hard Drive/RAID Errors

This error is sometimes challenging for even those with years of experience and expert administration skills. Whenever your hard drive fails, your server’s data is at risk.  If you see errors indicating a possible issue with one of your hard drives, a PSSC Labs support technician can assist you with checking the drive’s built in SMART (Self-Monitoring, Analysis and Reporting Technology) data, which will show the drive’s health and indicate whether or not it needs to be replaced.  It’s possible that the drive is healthy and there is another underlying issue which needs to be remedied, so please contact PSSC Labs should you encounter this error.

Machine is Reporting Memory/CPU Errors

PSSC Labs servers come equipped with large amounts of memory and speedy CPUs. When issues occur with these components, users will often see slow / degraded system performance accompanied by MCE (Machine Check Error) messages.  You can contact PSSC Labs and provide these messages to a PSSC Labs support technician who can assist you with determining which memory DIMM or CPU the error is coming from.  Memory can be easily replaced, whereas handling CPUs is an extremely delicate process — one that we recommend avoiding unless absolutely necessary. Often times MCE messages can be resolved by simply reseating or shuffling the memory around.

About PSSC Labs 

For technology powered visionaries with a passion for challenging the status quo, PSSC Labs is the answer for hand-crafted HPC and Big Data computing solutions that deliver relentless performance with the absolute lowest total cost of ownership.  

 We are true innovators offering high performance computing solutions to solve the world’s most demanding problems. For 25+ years, organizations of all sizes and from a variety of sectors rely on PSSC Labs’ computing systems. We are proud to support many departments within the United States government, Fortune 500 companies, as well as small and medium-sized businesses.  

All products are designed and built at the company’s headquarters in Lake Forest, California.

Government Agencies Leverage New Technology for Data Flow

By marketing@site-a.com

LAKE FOREST, CALIF., AUGUST 20, 2019 — PSSC Labs, a developer of custom High Performance Computing and Big Data computing solutions today shares how the CyberRax Data Flow Pipleline is helping various DOD organizations achieve their goals.

The Federal Government has spent decades collecting and storing huge data sets, covering everything from census and population records to crime reports and infrastructure data. The analytics retrieved from this data has assisted our government agencies in commissioning greater responsiveness and efficiencies at every level, making it truly invaluable to our society. Federal agencies simply can’t afford to suffer from computational roadblocks. Government servers need to be able to perform streaming analytics while also providing full control over their own data, especially when it comes to governance, security and performance.   

PSSC Labs, in partnership with Cloudera, is deploying the CyberRax Data Flow Pipeline as a truly turnkey system engineered specifically to meet the data transport needs and security requirements of Federal Agencies. The CyberRax Data Flow Pipeline comes complete with all necessary storage, network, operating system and application software including Cloudera Data Flow (CDF). CDF is built on top of Apache Nifi, a powerful and user-friendly data routing application originally developed by the National Security Agency (NSA). With CDF, users have the ability to create simple graphical models that direct data across multiple networks with all necessary data provenance required for the most secure environments.  

A Truly Turn-key Platform 

According to Alex Lesser, PSSC Labs Vice President, the real value of CyberRax Data Flow Pipeline is a production-ready platform out of the box. “What we see happening with the delivery of CyberRax Data Flow Pipeline is a significant reduction in time and cost, allowing federal agencies to see true value much sooner,” says Lesser. With 48 processor cores, 4 GB memory per core, 264TB RAW storage capacity and dual 10GigE network backplane, CyberRax Data Flow Pipeline can fulfill the most demanding environments. Federal agencies can leverage the engineering expertise of PSSC Labs high performance and high reliability platform. That’s why it only makes sense that government agencies including NASA, NIH, DOD and many more have partnered with PSSC Labs in an effort to quickly process, store, analyze and protect their massive amounts of data. 

The U.S. Air Force recently deployed several CyberRax Data Flow Pipelines for both their NIPR and SIPR networks. The goal of these deployments is to collect, transport, and distribute cybersecurity data from every Air Force installation worldwide on both NIPR and SIPR. This solution allows for data to be sent to multiple core locations and fed downstream to multiple cybersecurity tools for analysis. 

American Made

As an American manufacturer, PSSC Labs provides a level of security and trust that other foreign manufacturers simply lack. The CyberRax Data Flow Pipeline can be purchased via several contract vehicles including GSA, CHESS and NETCENTS. PSSC Labs is already working closely with Department of Defense agencies to help them achieve their operational mandate to ingest and direct data from disparate sources all with the required security provisions out of the box. We know that government servers need to be 100% operational at all time, without sacrificing speed, analytics, or security, which is why our team of engineers is available for support for the lifetime of your system.

About PSSC Labs 

For technology powered visionaries with a passion for challenging the status quo, PSSC Labs is the answer for hand-crafted HPC and Big Data computing solutions that deliver relentless performance with the absolute lowest total cost of ownership.  

 We are true innovators offering high performance computing solutions to solve the world’s most demanding problems. For 25+ years, organizations of all sizes and from a variety of sectors rely on PSSC Labs’ computing systems. We are proud to support many departments within the United States government, Fortune 500 companies, as well as small and medium-sized businesses.  

All products are designed and built at the company’s headquarters in Lake Forest, California. 

PSSC Labs Featured in IDG Connect: “Big Intelligence” is the real AI

By marketing@site-a.com

Everyone knows the scenario – after years of development and advancements, machines imbued with Artificial Intelligence somehow become self-aware without the knowledge of their human creators and end up destroying humanity as we know it. It’s a crazy premise, but if you listen to Tesla and SpaceX CEO Elon Musk and other futurists, it’s a possibility.

Science fiction movies have a habit of predicting doom and destruction, but the truth is we don’t have to worry about this dystopian world where AI becomes something man cannot control. True, there could be unpredictable social consequence much like those brought about by the rise of social media. But in terms of actual takeover and destruction, the odds are slim to none. Rather than fretting about killer robots, it’s time to realise that the AI revolution is actually the proliferation of “Big Intelligence” and its future is much more benign. Big Intelligence is where we are today in terms of automation, robotics and computing – and it bears little similarity to the sentient machines most people think of when they hear AI. What’s more, fear of AI shouldn’t hinder the legitimately useful work Big Intelligence can help complete.

AI is Big Intelligence in disguise

Many companies like to tout the AI capabilities of their offerings, but much of the literature you’ll find is nothing more than marketing gimmick. The technology being commercialised or used in research is not really AI, but simply better programming and faster data crunching fueled by advancements in hardware. While futurists like to envision the all-knowing, self-aware machine, Artificial Intelligence or machine learning is not a substitute for human intelligence, and even the most advanced machines are far from substituting a human brain.

With the convergence of large datasets with faster computers and better code to process the data, Big Intelligence has progressed over the past decade due to real technology advancements that allow us to collect and process data at an ever increasing rate, and it’s what powers most AI platforms today. But the key here is not some technology that can replace human intelligence – rather, AI as it currently stands is a set of tools that merely help us better process, interpret, and understand the mass amounts of data companies gather from actual thinking humans. It can help businesses better predict their own and their customers’ needs, to better optimise and even conserve resources.

Hardware is critical to Big Intelligence

The only reason we are discussing the possibility of AI is due to recent advancements in computing performance. If the hardware could not process data in near real-time, things like self-driving cars, automated logistics centres, operating systems that are virtual and learn through association would still be figments of our imagination.  However, ultimately these so-called AI data still rely on constant data input from humans, without which they could not function.

Artificial Intelligence is a misnomer. Most of what we consider AI is really a high-performance computer (HPC) crunching a massive amount of data, which does not have intelligence the way a human brain does to conceptualise, reach its own conclusions and think for itself. What it can do is improve computing processes through automation, but the end result of most automated processes still requires human supervision.

Computers can now process data in real time, but all that processing power is useless if people feed the machine bad data from the get go or don’t know what to do with all that information and analysis once they have it. Cognitive solutions that leverage AI can provide explanations, recommendations, and inform what future actions or outcomes might be required via their predictive nature, but it’s still a human who is feeding the beast.  We are still the “intelligence” behind AI – the artificial part is being able to crunch data at a scale and time-span humans can’t achieve.

The promise of AI has been around a long time, but never went anywhere because hardware could not sustain that much data analysis. No one could capitalise on the concept. That is not the case today. Hardware has advanced to such a degree that for each new automation concept there is a company that builds the hardware necessary to realise the idea.  Super-fast multi-core processors or massive storage devices are tomorrow’s recycling candidate. Cloud computing, virtualisation, faster processors – all make up the core technology of what we call AI.

In addition, the amount of data companies now churn out is almost unfathomable. Almost everything we touch is sending data to someone from smartphones, to the internet, to online processes, to set-top boxes, and on and on. Technology is everywhere and every bit of it is a data source. At the convergence of all this Big Data and mock AI technology is nothing that resembles actual sentience – it’s simply Big Intelligence. And we’ll continue to see it improve as this convergence of better programming through High Power Computing and Big Data as companies traditional working with one or the other begin to bring the two together in ever more creative applications.

We are living in a truly exciting time of data and high-performance computing. We now have the ability to take advantage of data at scale and analyse that data in real time. But we shouldn’t let the misnomer of AI scare us away from this pursuit. Let’s call Big Intelligence what it is and enjoy the power more data and advanced hardware has bestowed on the human race.

Source: http://www.idgconnect.com/blog-abstract/28283/-big-intelligence-real-ai