PSSC Labs Featured on The Next Web: Why are companies choosing on-premise HPC over cloud?

By marketing@site-a.com

The concept of High-performance Computing (HPC) in the cloud has taken a massive leap forward these past years. While the idea of using High-Performance Computing (HPC) services like storage, servers, databases, networking and software applications etc over the cloud isn’t new, what is new is the speed and commitment companies have placed in the cloud. However, with the expansion of business, some companies look for on-premise options. Why? Because the benefits of High-Performance Computing (HPC) in the cloud are manifold but it is also countered by some drawbacks that are leading organizations to look at other options. Let’s have a contrasting discussion on both!

Supercomputing in the cloud
Cloud computing is explained by PCMag in the simplest way possible as storing, managing and accessing data and programs over the internet instead of on machine. The best examples we are all familiar with include Google Drive, Google Docs, Microsoft OneDrive and Dropbox as cloud storage applications and more sophisticated cloud service offered is High-performance Cloud Computing that include Amazon Web Services, Microsoft Azure, Google, and IBM.

Cloud computing can provide multiple features to an average user, like instant availability of resources, availability of large capacity for storage and processing, flexibility at the application level and a bare minimum level of performance guaranteed by the provider. However, the user for HPC generally deviates from all these features and presents a tailored requirement for their specific application. A hardware fine-tuned to the needs of the application is the dream of such users; they are not here to run generic applications.

They often try to get rid of the OS formalities and start talking to the hardware directly. The cloud OS “nanny” might not allow you to get too direct with her baby, whereas, HPC applications need to bypass the OS kernel a lot.

When is it ideal to get cloud based HPC?
Most applications are very sensitive about the networks interconnect; the data needs to flow at speeds to match the high-performance demands. A virtual cluster is limited to the rules defined by the kernel and many high-performance network loads need to manage the connection and data transfers ‘on the wire’, which is very hard to visualize on a virtual scheme. Along with it comes the requirement of design specific storage system. A strong I/O system is next on your requirement list, without that you are likely facing bottlenecks, backlogs and unnecessary queuing most of the times.

A cloud-based HPC is a very good bargain while working with rudimentary tasks, maybe for a startup, small business ventures or an on-demand test facility with a limited influx of tasks. But deeper and more elaborate discussion would go into the decision of using cloud-based HPC against on-premises for anything bigger than that.

On-premise Supercomputers
When you think of supercomputers, the first image that streams through your mind is one of massive rows of mainframe computers filling an entire room with lots of noise and huge cooling pipes circling around it.

That was true at least half a century ago. Today, supercomputers (referred to as High-Performance Computers) can perform all your high demand computing tasks running advanced applications and manage large data sets with advanced network management tools in much more compact servers or clusters. Clusters of HPCs share the workload by dividing the tasks into parts and feeding them to the parallel processing units of a supercomputer (as opposed to serial processing of a normal computer).

Advances in technology mean that today’s supercomputers come in compact designs (in a sleek 1U and 2U size casing) with less maintenance and resource demands. For example, the PowerServe HPC by PSSC Labs 16 to 72 total cores of Intel® Xeon processors with as much as 1024 GB of high-performance memory packed in a 1/2U blade chassis, versatile network connectivity options and supporting all the latest operating systems. Even better, their unique design means a 90% energy efficient power supply. So having an on-premises supercomputer is not so ‘super’ difficult at all.

Why are companies choosing on-premise HPC over the cloud?
Amir Michael, a former hardware engineer at Google and former hardware and data center engineer at Facebook, founder, and CEO of Coogan, says “Surprisingly, a lot of people are thinking about off-boarding from the cloud and trying to figure out when the right time might be to do that. Other customers are pretty big in co-location and they are wondering if they should build their own data centers.”

This calls for a good night long debate. Companies today see on-premise HPCs as a liability because of the high purchasing costs and associated maintenance costs. Maintenance also comes with the potential need to hire more personnel to maintain the infrastructure. So naturally, outsourcing this to AWS, Google, Microsoft, etc. seem like the best way to go about it. However, this step can be shortsighted and vitally dangerous. In addition, putting your entire business (data, analytics and most importantly intellectual property) on the cloud means you are giving up control of the lifeblood of your business.

Reasons why on-premise HPC is taking the lead
Let us consider some of our own parameters and see how that puts our computing needs into perspective:

Performance: 
With your own HPC infrastructure (whether just a rack server or a cluster), you achieve a much better performance per dollar per hour as compared to any other generic server. Since the hardware you have is to meet your specific application requirement, you are achieving the optimum level on the cost-performance chart. By going with a dedicated on-premise option you can design your hardware that complies with your exact needs.

Cost:
The main reason cloud computing has soared in terms of adoption is the belief that outsourcing your computing needs to the cloud is cheaper than doing it yourself. The answer isn’t always so clear-cut. When scoping out Total Cost of Ownership (TCO), factors may arise that aren’t in the original calculation which leads to a higher long-term TCO for trying to do HPC in the cloud. As many companies are realizing who are moving off AWS to go in-house, the promise of lower cost through the cloud isn’t quite so clear-cut.

Access to Data:
Continuing on the last note, every time you want to access your own data (which you are keeping at Azure server for as low as $0.02), you have to pay a price to retrieve that data. On-premises HPC grants access to your data any time you need. A popular solution for startups is to use NAS (Network Attached Storage) solutions by vendors such as Seagate. But for more complex computing projects, you may need a scalable block and object storage platform like the Surestore by PSSC Labs.

Security
AWS and others have made strides here but it was not long ago that a simple keystroke error brought down nearly 30% of the websites on the east coast. Companies need to evaluate putting their livelihood into someone else’s hands versus the peace of mind of having critical HPC functions close at hand and under your control.

Source: https://thenextweb.com/guests/choosing-on-premise-hpc-over-cloud

PSSC Labs Introduces Affordable, High Performance All Flash Storage on Unique Enterprise Server Designed Specifically for Hadoop

By marketing@site-a.com

Custom, high performance server solution now with storage option from Micron’s powerful SSD line

LAKE FOREST, Calif., May 1, 2017 /PRNewswire/ — PSSC Labs, a developer of custom HPC and Big Data computing solutions, today announced it now offers the option of complete flash storage for its CloudOOP 12000 enterprise server, built specifically for Hadoop. The faster storage option means PSSC Labs can now offer increased data analytics speeds and improved performance, making the CloudOOP 12000 the ideal enterprise Hadoop solution.

The CloudOOP 12000 now supports all flash storage that can replace the disk drives shown here, maintaining the CloudOOP 12000’s innovative design while delivering even faster analytics.

The CloudOOP 12000 now supports the new Micron 5100 series Solid State SATAIII (SSDs), with up to 14 x Micron SSDs (2 drives dedicated for operating system and 12 drives for data storage). The all flash storage option provides durability, reliability, and faster performance at a price point that makes the technology affordable to more enterprise users. The Micron SSD’s offer faster speeds that traditional non-flash storage, and when combined with the CloudOOP 12000’s made for Hadoop server, can achieves near real time performance for data analytics.

Using the largest capacity Micron SSDs, the CloudOOP 12000 can support up to 96 TBs of storage and is tailored to meet the needs of read-intensive video streaming, latency-sensitive transactional databases and write-intensive logging applications. It also makes an excellent platform for edge computing with the ability to consume large amounts of data very quickly from IOT devices and sensors.

The CloudOOP 12000 is the only server specifically designed for Hadoop, Kafka, Big Data and IOT. It offers 2x the density and up to 35% lower power draw than traditional manufacturers as well as a near 50% increase in data throughput performance. Reducing power draw means a lower data center footprint and significantly reducing your total cost of ownership with an over 90% efficiency rating.

“The CloudOOP1200 is PSSC Lab’s unique platform for enterprise users using applications like Hadoop, Spark, Kafka Streaming. PSSC Labs has already successfully deployed over 100 PBytes for Hadoop using the CloudOOP 12000 platform, and after stringent review of the SSD options on the market, we’ve certified of Micron’s new 5100 series of SSDs, which will allow us to offer our customers a high capacity, high performance, durable system with the absolute lowest cost of ownership, “said Alex Lesser, Vice President of PSSC Labs.

Micron 5100 Series SSD Features include:

High Capacity
Unique range of solutions with up to 8TB of storage in a 2.5-inch form factor and 2TB in an M.2

High Performance
Three models optimized for varying workloads with consistent, steady state random writes at 74,000 IOPS.

Secure Encryption
Built-in AES-256-bit encryption and TCG Enterprise protection with FIPS 140-2 validation – available on the 5100 MAX.

Greater Flexibility
Micron’s FlexPro firmware architecture can be used to actively tune capacity to optimize drive performance and endurance

Best Reliability
Unmatched 99.999%2 quality of service (QoS) compared to spinning media. MTTF of 2 million device hours

CloudOOP 12000 Features Include:

High Processing Power
The CloudOOP 12000 supports up to 2 x Intel Xeon E5 Series processors & up to 256GB high performance memory – get higher performance and reduced computing time

Direct Connect IO technology
Unique design gives each hard drive its own independent path to the motherboard – removing unnecessary components that restrict data pathways and improving data ingestion & IO rates.

Connectivity Options
GigE, 10GigE, 40GigE and Infiniband network connectivity options available. Dual GigE network bandwidth comes standard, with addition network adapters from Intel, Mellanox, Solarflare and others available.

Operating System Compatibility
Supports Microsoft Windows, Red Hat, CentOS, Ubuntu & most other Linux distributions.
All CloudOOP 12000 server configurations service and support from PSSC Lab’s US based, expert in house engineers. Prices for a custom CloudOOP 12000 server start at $5000.

For more information see https://pssclabs.losangeles.dev.buckupstudio.com/products/big-data-servers/.

Source: http://www.prnewswire.com/news-releases/pssc-labs-introduces-affordable-high-performance-all-flash-storage-on-unique-enterprise-server-designed-specifically-for-hadoop-300448604.html

Whitepaper: Loading a Time Series Database at 100 Million Points Per Second.

By marketing@site-a.com

There are many use cases for time series data, and they usually require handling a decent data ingest rate. Rates of more than 10,000 points per second are common and rates of 1 million points per second are not quite as common, but are certainly not unheard of.

Many of these systems are used to monitor critical infrastructure, especially at the higher data rates, where failure of the monitoring systems can lead directly to disastrous failures of the actual system.

It’s important to determine that your time series database will actually work at full loads, so naturally, many organizations will test it at full load. To do this correctly, you have to first fill the database with typically a year or 3 of data, which is where you will suddenly find yourself with a much higher data rate requirement. It makes no sense to test a system with 3 years of data if it takes 3 years to fill the database in preparation for the test.

Thus, if you have a production data rate of 100,000 points per second, you will need to ingest test data at a rate of 100 million points per second if you want to load 3 years of data in less than 2 days. Even a production rate of 10,000 points per second will require several hours to load test data.

Similarly, if you are starting a new database, you probably will want to load significant amounts of historical data, which could be just as stressful as loading the database for testing.

How did we do this?

There are three major components that should be considered:

  1. Loading one point at a time is very slow; batching must be implemented in order to accomplish faster ingestion rates.
  2. The data store needs to be highly performant in order to consistently handle such massive amounts of data.
  3. Well-configured hardware is critical to achieving the highest performance.

First, OpenTSDB was used for this test because it is a well documented time series database and it uses the HBase API for its data store. OpenTSDB also uses the non-relational nature of the HBase API to strong advantage by storing data in a hybrid wide data schema. In our tests, we used one year of data. Sampling one sensor every second for one year generates about 31.5 million points. In OpenTSDB’s storage format, this translates to about 120MB of data per sensor year of data.

Second, MapR-DB was used for this test instead of HBase itself. MapR-DB offers many benefits over HBase, while maintaining the virtues of the HBase API and the idea of data being sorted according to primary key. MapR-DB provides operational benefits such as no compaction delays and automated region splits that do not impact the performance of the database. The tables in MapR-DB can also be isolated to certain machines in a cluster by utilizing the topology feature of MapR. This allowed the testing to easily be scaled to measure the performance on any number of nodes in the cluster. The final differentiator is that MapR-DB is just plain fast, due primarily to the fact that it is tightly integrated into the MapR file system itself, rather than being layered on top of a distributed file system that is layered on top of a conventional file system.

Third, the hardware that was used was fast. The test cluster was designed by PSSC Labs specifically for use with Hadoop systems. The CloudOOP 12000 platform offers a unique, highly optimized system design with a direct data path for disk I/O for up to 12 SATA/SAS devices in a 1u rackmount chassis. This hardware explicitly lacks a raid controller and a backplane, as its intended use is for Hadoop and would otherwise be unused. Here are the specifications for this test setup:

  • Dual Intel E5-2650 v2 Processor (16 total physical cores, 32 hyper-threads)
  • 128GB DDR3 1600
  • 1 – 120GB SSD for the operating system
  • 11 – Western Digital 1TB 7200 RPM 6GB/s
  • Solarflare SFN5152 10GigE network adapter
  • CentOS 6.x
  • Estimated power draw: 240 watt idle / 325 watt 100% load

How were these numbers achieved?

Code was created by modifying portions of OpenTSDB to allow bulk importing of data. This code has been published in the MapR App Gallery and is also available via github at https://github.com/mapr-demos/opentsdb. Fundamentally what is happening is that OpenTSDB stores all points that are for a given hour into a single row in MapR-DB. During normal operations, OpenTSDB inserts each point into the row in a separate column. Once an hour, the entire row is read, columns are collected into a blob and the row is written back to the database. This results in a read and a write per point. In our bulk loading code we directly create the blob for an entire hour of data (3,600 points in our test) and insert that blob into the database in a single put. This single put takes the place of over 7,000 database operations per sensor-hour of data.

We used two edge nodes, both running code to generate and load test data. While these edge nodes were generating data, we observed near 100% CPU utilization on the edge nodes. These edge nodes were not quite hitting their outbound bandwidth limit, as they couldn’t generate data any faster than the CPU bottleneck would allow. To achieve maximum performance, we used eight separate Java virtual machines, each running two threads actively generating the data. Other configurations with more or fewer JVMs and more or fewer threads per JVM gave equal or inferior results.

We limited the output data to only four nodes out of a ten node cluster. Our original plan was to start with four nodes and increase the nodes involved in the test until we achieved our goal of more than 100 million points per second. This test generated random integers that emulated 128 sensors sampled at one second intervals for one year. This yielded over four billion points totaling about 15GB on disk storage. While this data size is less than the memory available on the test nodes, writes to the database are all fully persisted to disk during the test. The ingest rate was approximately 110 million points per second and the data load completed in about 36 seconds. The MapR-DB table was set to the default of 4GB regions, which meant that multiple region splits would occur during the tests. The data replication for the table was set to three, meaning every record was duplicated from the first server to two other servers.

At peak loading, we observed that the two edge nodes together were pushing approximately 1.5GB/s of total data from the edge nodes to the cluster nodes. Inbound data volumes on individual cluster nodes were variable, with peaks near 1.3GB/s. Impressively, the 7200 RPM disks were able to keep up with the network speed and were able to persist data with no issues. Data generation was limited by CPU capacity on the edge nodes. Ingest volume was limited by network bandwidth on the cluster side. It was clear that our loading did not drive the disk subsystems to capacity. Adding more cluster nodes will distribute the network load within the cluster, but we expect that the edge node count is the primary limit on ingestion in the current test configuration.

Could it get any faster?

The storage format of OpenTSDB is reasonable, but could be improved significantly. For instance, time values are stored in the column name for the blob of data. Using a single column name and a compressed binary representation of the data would allow substantial compression of the data. One example of this compression is in the times themselves. With delta coding, the total storage for each time in a blob format could be reduced to less than 1% of the current size. The flags which indicate the time resolution of the timestamps and whether the data is floating point could be repeated only once. For sensor data, resolution is typically limited so that dictionary encodings would likely result in up to 10x compression.

In terms of hardware, the networking bandwidth could be increased, perhaps utilizing dual 10GbE NICs or even higher-end networking gear. The MapR distribution supports transparent application level multiplexing of multiple network interfaces, allowing over 2GB/s of data to be transferred over the network. There is also a possibility that performance would improve with faster disks. These changes would increase the cost of the servers, so it would be worthwhile to understand the cost-benefit of the actual gains.

Performance could also be trivially increased by scaling the system out so that the main data table is spread across more than four nodes.