PSSC Labs Featured in Cloud Tech News: Four industries where on-premises infrastructure beats the cloud

By marketing@site-a.com

Cloud computing has fundamentally transformed information technology by offering enterprises an allegedly cheaper, more flexible, and relatively maintenance-free alternative to purchasing their own IT infrastructure. However, despite the mad rush to migrate to the cloud, cloud solutions are not always the best or least expensive choice, particularly in industries that work with very large, complex data sets, perform many intricate mathematical calculations, possess invaluable digital intellectual property, or are subject to certain compliance standards.

Let’s take a look at four industries where in-house IT infrastructure still beats the cloud.

Cybersecurity

In the cybersecurity field, the ability to process extremely large data sets as quickly as possible is crucial, especially as cyber attacks shift away from “lone wolf” one-off hacks and towards highly organized, sophisticated operations carried out by well-trained cyber criminals. IBM estimates that the average organization encounters an average of 200,000 security event alerts each day. An enormous amount of computing power is required to run the SIEM systems that not only detect those anomalies, but also analyze them, separate the false positives from the possible attacks, and deliver actionable information to security analysts – way more power than even the most robust cloud solution could possibly deliver.

In addition to latency problems, the amount of bandwidth consumed would become very expensive, very quickly, making on-premises equipment the most cost-effective solution over time. On-premises cybersecurity hardware also prevents chain-reaction situations where client organizations end up getting hacked because the cloud vendor that their cybersecurity provider was using did.

Adtech

The internet has transformed the way in which consumers and businesses shop. Long buying cycles have been replaced by “just-in-time,” last-minute purchasing decisions, and advertising-weary prospects are using ad blockers, email filters, and DVR fast-forward functions to tune out traditional advertising. To reach these elusive prospects, companies are turning to ad tech, which employs extensive market research and big data analytics to deliver highly targeted marketing messages to qualified prospects at the precise time that they are ready to buy.

The complex data sets and high-level analytics that power the ad tech industry require computing power that is already beyond what any could solution could offer. As the industry matures, ad tech firms will require even more power and more storage space. On-premises equipment can be scaled much more quickly than cloud solutions, and latency problems are avoided.

The ad tech industry also faces intellectual property issues. As Amazon, Google, and other cloud providers enter the market research and ad tech spaces themselves, questions arise as to the safety of digital intellectual property stored on a competing company’s cloud service. Dropbox listed Amazon’s move into the file-sharing space as one of the reasons why it decided to ditch AWS for its own equipment.

Life sciences

The life sciences industry is grappling with a “data avalanche” of clinical results, disease states, scientific studies, and individual patient data, which results in stratospheric cloud bills and latency problems. Because researchers work on limited budgets, simply storing all of this data on the cloud could deplete a project’s funding – and that’s before anything is actually done with it.

Cyber security and compliance issues also come into play due to the sensitive nature of this data. Storing patient data on the cloud may result in an organization running afoul of HIPAA and other privacy regulations if the cloud provider gets hacked, even if the hack turns out to be the provider’s fault.

Finally, on-premises hardware, unlike cloud solutions, keeps running even if the internet is down, which makes on-site equipment a must for scientists who are performing research in remote areas where internet access is spotty.

Design and engineering

Much like cyber security, ad tech, and the life sciences, design and engineering involves performing intricate calculations on very large, complex data sets, as well as running memory-intensive software and generating terabytes of new data every day. A cloud solution would be both wildly expensive and far too slow. In particular, cloud servers offer very poor interdomain communication; a physical server equipped with a high-performance communication architecture such as Intel’s Omni-Path is less expensive than a cloud solution is less expensive and offers low communication latency, low power consumption, and a high throughput.

Design and engineering companies also face a constant threat from digital intellectual property theft; everything from new product prototypes to R&D data could be targeted by competitors, cyber extortionists, or even foreign governments. Digital IP is simply too sensitive to be stored in the cloud, especially since hackers are increasingly targeting cloud providers as the industry grows.

Beyond cloud-first hype, a balanced approach

Despite its drawbacks, cloud computing does have a place in many organizations’ IT ecosystems. Many companies use on-premises equipment to handle standard functions and store highly sensitive data and employ cloud solutions when they require additional capacity or to store less-sensitive data.

Instead of migrating to the cloud because it’s trendy, and “everyone” says it costs less and offers more flexibility than in-house equipment, enterprises should take a step back, consider their individual computing needs, and perform an objective cost analysis.

Source: https://www.cloudcomputing-news.net/news/2017/dec/04/four-industries-where-premises-infrastructure-beats-cloud/

PSSC Labs Named a 10 Best Hadoop Solution Provider Companies to Watch

By marketing@site-a.com

Providers of custom turnkey platforms built specifically for Hadoop recognized in leading industry magazine

LAKE FOREST, Calif., Oct 18, 2017 /PRWEB/ — PSSC Labs, a developer of custom HPC and Big Data computing solutions, today announced it has been named to Insights Success Magazine’s 10 Best Hadoop Solution Provider Companies to Watch. Providing industry leading server platforms engineered specifically for Hadoop, PSSC Labs is a leader in providing turn-key solutions for enterprise Big Data needs.

Hadoop is a constantly evolving and expanding ecosystem that is disrupting the traditional storage and analytics platforms with more flexibility, speed, and reliability – creating an ideal, low cost, and scalable solution for today’s enterprise market. PSSC Labs’ exceptional record of offering the most powerful Hadoop solutions to its customers around the world has earned the company a place on this year’s Best Hadoop Solution Providers list. Editors at Insight Success Magazine analyze companies large and small to find the most unique and disruptive Hadoop solution providers on the market, choosing PSSC Labs as one of its standouts. To view the complete profile visit http://www.insightssuccess.com/pssc-labs-delivering-hand-crafted-hpc-extraordinary-big-data-computing-solutions/.

Designed for Hadoop

PSSC Labs has deployed over 100 petabytes of Hadoop infrastructure, including small POC clusters to production environments exceeding 10 PB. PSSC Labs’ CloudOOP Big Data Server Line is ideal for organizations looking for Hadoop solutions, and ideal for numerous industries including Ad Tech, Healthcare, Telecom, Media & Entertainment, and Cybersecurity. Designed specifically for Hadoop, Kafka, Big Data and edge computing with IOT devices and sensors, the CloudOOP line is engineered for faster data ingestion and processing.

Proven compatible with leading data platforms including Hortonworks, Cloudera and MapR, the line offers enterprises a platform that will significantly lower both capital expenditure and operating expenses while achieving higher data throughput performance. Utilizing PSSC Labs’ specialized server design and unique, customized solutions, the CloudOOP line delivers high performance, high density, flexibility and scalability while reducing overall footprint and power usage. Offering 2x the density, 35% reduction in power use and up to 50% increase in data throughput, the CloudOOP line allows companies to scale their infrastructure while reducing the need for additional network, rack and power components – reducing total cost of ownership.

All CloudOOP server configurations receive service and support from PSSC Labs’ US-based, expert in-house engineers. For more information and configuration options see https://www.site-a.com/servers/big-data-servers/

Source: http://www.prweb.com/releases/2017/10/prweb14808617.htm

PSSC Labs Featured in IT Pro Portal: 7 hidden AWS Costs That Could Be Killing Your Budget

By marketing@site-a.com

Is your organization paying what it should for hosting with AWS?

The AWS Elastic Compute Cloud (EC2) service has many advantages, including easy scalability, pay-for-what-you-use, as-you-go pricing, and an enormous array of options and upgrades – so many that your AWS bill may become quite complicated. Have you been suffering from sticker shock but have no idea which of the literally thousands of line items on your invoice are the culprits? Here are seven hidden AWS costs that could be breaking the bank. 

1. Unused Instances 

Among the biggest contributors to inflated AWS bills are unused or underutilized EC2 instances, which result in your organization paying for resources it is not using. Be sure to terminate all instances as soon as you are finished using them; note that this applies separately, to each world region. Also make sure to monitor your EC2 usage data for low CPU usage, bandwidth, and I/O, which are red flags that may indicate underutilized servers that could be shut down. 

2. Unneeded Orphaned Snapshots 

Terminating an unused or underutilized EC2 instance isn’t enough. Even though the attached EBS volumes are automatically deleted along with the instance, your snapshots will remain stored on Amazon Simple Storage Service (S3), and you will continue to be charged monthly for them. Deleting these orphaned snapshots will save as much money as deleting the original EBS volume, so unless you’re certain you’ll need them again to create future EBS volumes, make sure to get rid of them. 

3. Unattached/Unused EBS Volumes 

It’s good practice to delete EBS volumes, except for root volumes, that are not attached to EC2 instances. Not only do these unattached volumes add charges to your AWS bill – whether they are being used or not – but they also pose cyber security risks to any sensitive data stored on them. Even when an EBS volume is attached to an EC2 instance, it’s billed for separately, so make sure to delete volumes you no longer need. Just make sure to back up the data first; once the EBS volume is deleted, the data is lost. 

4. Underutilizing Reserved Instances 

Many AWS users pay as they go and never consider purchasing reserved instances (RIs), which are pre-booked resources and capacity for a one- or three-year term. Because you are committing to pay for all of the hours during your term, your hourly rate is deeply discounted. RIs can save your organization a lot of money – if you actually use all of the time you’ve bought. Calculating your usage that far in advance can be difficult, and if your needs are lower than you anticipated, RIs can be a money pit. If you’ve bought RIs that you have no use for, consider selling them on the AWS Marketplace. 

5. Data Transfer Costs 

Most of the time – but not always – transferring data into EC2 is free, but transferring data out will always cost you. How much you will be charged depends on how much data is being transferred and where it is going, and these costs vary by Region. Moving data across services within the same Region is usually less expensive than moving data across services outside your Region, and some Regions have higher costs than others. To minimize data transfer costs, you must pick the least expensive route for your data to flow through, depending on what it is and where it is going; because of the wide variations in prices, this is easier said than done. 

6. Unused Elastic IP (EIP) Addresses 

EIPs are static different EC2 instances. They allow users to mask an instance or software failure by rapidly remapping the address to another instance in their account. AWS users are allotted one free EIP address with each running EC2 instance, but they are charged hourly if they attach additional EIPs to the same instance. Additionally, users are charged for any EIPs that are not associated with a running instance. As with instances, orphan snapshots, and EBS volumes, you should monitor your account for EIPs you are no longer using. 

7. Unused Elastic Load Balancers (ELBs) 

ELBs, which are commonly put in front of your web servers, automatically distribute incoming application traffic, scale resources to meet traffic demands, and are designed to keep a minimum number of EC2 instances running. You are charged monthly for each ELB, whether you’re using it or not, and per GB transferred. If any of your ELBs are not attached to back-end instances, consider registering instances or deleting them. In a similar vein, if an ELB is not attached to any healthy backend instances, consider troubleshooting the configuration or deleting it. Additionally, before you can terminate an EC2 instance, you must delete any ELBs associated with it. 

Other Hidden Costs 

Other potential “gotcha’s” that could be inflating your AWS bill include unused services started in AWS OpsWorks, unhealthy instances, and fees for excessive API calls. There are also numerous indirect costs associated with AWS and other cloud solutions in the form of performance, reliability, and cyber security problems. Misconfigured AWS servers were at fault for the recent data breaches at business associates of Verizon, the Republican National Committee, and private security firm TigerSwan. In February, numerous large websites were knocked offline due to an error by an employee at AWS, and the tech community recently expressed grave concerns about widespread chaos if AWS were to have another, larger failure, particularly since so many financial institutions rely on it. 

The Cloud Isn’t Always Cheaper 

Despite sticker shock and concerns about these hidden and indirect costs, many organizations continue to grumble and pay their AWS bill due to the misconception that cloud computing is always cheaper and more efficient than purchasing their own IT infrastructure. This is a myth. In many cases, an organization’s monthly AWS bill alone costs more than an in-house solution would. If your organization processes large amounts of data, it would probably be more cost-effective to purchase and maintain your own infrastructure. 

It’s not always necessary or beneficial to abandon the cloud completely. Many organizations would greatly benefit from a hybrid approach, where they use their own infrastructure for certain tasks and utilize cloud solutions when they need additional capacity. 

Don’t feel like you’re locked into paying AWS or another cloud provider forever. If you can’t seem to get your AWS bill down to a reasonable level, purchasing your own equipment is worth looking into. 

Source: http://www.itproportal.com/features/7-hidden-aws-costs-that-could-be-killing-your-budget/

PSSC Labs Featured on The Next Web: Why are companies choosing on-premise HPC over cloud?

By marketing@site-a.com

The concept of High-performance Computing (HPC) in the cloud has taken a massive leap forward these past years. While the idea of using High-Performance Computing (HPC) services like storage, servers, databases, networking and software applications etc over the cloud isn’t new, what is new is the speed and commitment companies have placed in the cloud. However, with the expansion of business, some companies look for on-premise options. Why? Because the benefits of High-Performance Computing (HPC) in the cloud are manifold but it is also countered by some drawbacks that are leading organizations to look at other options. Let’s have a contrasting discussion on both!

Supercomputing in the cloud
Cloud computing is explained by PCMag in the simplest way possible as storing, managing and accessing data and programs over the internet instead of on machine. The best examples we are all familiar with include Google Drive, Google Docs, Microsoft OneDrive and Dropbox as cloud storage applications and more sophisticated cloud service offered is High-performance Cloud Computing that include Amazon Web Services, Microsoft Azure, Google, and IBM.

Cloud computing can provide multiple features to an average user, like instant availability of resources, availability of large capacity for storage and processing, flexibility at the application level and a bare minimum level of performance guaranteed by the provider. However, the user for HPC generally deviates from all these features and presents a tailored requirement for their specific application. A hardware fine-tuned to the needs of the application is the dream of such users; they are not here to run generic applications.

They often try to get rid of the OS formalities and start talking to the hardware directly. The cloud OS “nanny” might not allow you to get too direct with her baby, whereas, HPC applications need to bypass the OS kernel a lot.

When is it ideal to get cloud based HPC?
Most applications are very sensitive about the networks interconnect; the data needs to flow at speeds to match the high-performance demands. A virtual cluster is limited to the rules defined by the kernel and many high-performance network loads need to manage the connection and data transfers ‘on the wire’, which is very hard to visualize on a virtual scheme. Along with it comes the requirement of design specific storage system. A strong I/O system is next on your requirement list, without that you are likely facing bottlenecks, backlogs and unnecessary queuing most of the times.

A cloud-based HPC is a very good bargain while working with rudimentary tasks, maybe for a startup, small business ventures or an on-demand test facility with a limited influx of tasks. But deeper and more elaborate discussion would go into the decision of using cloud-based HPC against on-premises for anything bigger than that.

On-premise Supercomputers
When you think of supercomputers, the first image that streams through your mind is one of massive rows of mainframe computers filling an entire room with lots of noise and huge cooling pipes circling around it.

That was true at least half a century ago. Today, supercomputers (referred to as High-Performance Computers) can perform all your high demand computing tasks running advanced applications and manage large data sets with advanced network management tools in much more compact servers or clusters. Clusters of HPCs share the workload by dividing the tasks into parts and feeding them to the parallel processing units of a supercomputer (as opposed to serial processing of a normal computer).

Advances in technology mean that today’s supercomputers come in compact designs (in a sleek 1U and 2U size casing) with less maintenance and resource demands. For example, the PowerServe HPC by PSSC Labs 16 to 72 total cores of Intel® Xeon processors with as much as 1024 GB of high-performance memory packed in a 1/2U blade chassis, versatile network connectivity options and supporting all the latest operating systems. Even better, their unique design means a 90% energy efficient power supply. So having an on-premises supercomputer is not so ‘super’ difficult at all.

Why are companies choosing on-premise HPC over the cloud?
Amir Michael, a former hardware engineer at Google and former hardware and data center engineer at Facebook, founder, and CEO of Coogan, says “Surprisingly, a lot of people are thinking about off-boarding from the cloud and trying to figure out when the right time might be to do that. Other customers are pretty big in co-location and they are wondering if they should build their own data centers.”

This calls for a good night long debate. Companies today see on-premise HPCs as a liability because of the high purchasing costs and associated maintenance costs. Maintenance also comes with the potential need to hire more personnel to maintain the infrastructure. So naturally, outsourcing this to AWS, Google, Microsoft, etc. seem like the best way to go about it. However, this step can be shortsighted and vitally dangerous. In addition, putting your entire business (data, analytics and most importantly intellectual property) on the cloud means you are giving up control of the lifeblood of your business.

Reasons why on-premise HPC is taking the lead
Let us consider some of our own parameters and see how that puts our computing needs into perspective:

Performance: 
With your own HPC infrastructure (whether just a rack server or a cluster), you achieve a much better performance per dollar per hour as compared to any other generic server. Since the hardware you have is to meet your specific application requirement, you are achieving the optimum level on the cost-performance chart. By going with a dedicated on-premise option you can design your hardware that complies with your exact needs.

Cost:
The main reason cloud computing has soared in terms of adoption is the belief that outsourcing your computing needs to the cloud is cheaper than doing it yourself. The answer isn’t always so clear-cut. When scoping out Total Cost of Ownership (TCO), factors may arise that aren’t in the original calculation which leads to a higher long-term TCO for trying to do HPC in the cloud. As many companies are realizing who are moving off AWS to go in-house, the promise of lower cost through the cloud isn’t quite so clear-cut.

Access to Data:
Continuing on the last note, every time you want to access your own data (which you are keeping at Azure server for as low as $0.02), you have to pay a price to retrieve that data. On-premises HPC grants access to your data any time you need. A popular solution for startups is to use NAS (Network Attached Storage) solutions by vendors such as Seagate. But for more complex computing projects, you may need a scalable block and object storage platform like the Surestore by PSSC Labs.

Security
AWS and others have made strides here but it was not long ago that a simple keystroke error brought down nearly 30% of the websites on the east coast. Companies need to evaluate putting their livelihood into someone else’s hands versus the peace of mind of having critical HPC functions close at hand and under your control.

Source: https://thenextweb.com/guests/choosing-on-premise-hpc-over-cloud

PSSC Labs Introduces Affordable, High Performance All Flash Storage on Unique Enterprise Server Designed Specifically for Hadoop

By marketing@site-a.com

Custom, high performance server solution now with storage option from Micron’s powerful SSD line

LAKE FOREST, Calif., May 1, 2017 /PRNewswire/ — PSSC Labs, a developer of custom HPC and Big Data computing solutions, today announced it now offers the option of complete flash storage for its CloudOOP 12000 enterprise server, built specifically for Hadoop. The faster storage option means PSSC Labs can now offer increased data analytics speeds and improved performance, making the CloudOOP 12000 the ideal enterprise Hadoop solution.

The CloudOOP 12000 now supports all flash storage that can replace the disk drives shown here, maintaining the CloudOOP 12000’s innovative design while delivering even faster analytics.

The CloudOOP 12000 now supports the new Micron 5100 series Solid State SATAIII (SSDs), with up to 14 x Micron SSDs (2 drives dedicated for operating system and 12 drives for data storage). The all flash storage option provides durability, reliability, and faster performance at a price point that makes the technology affordable to more enterprise users. The Micron SSD’s offer faster speeds that traditional non-flash storage, and when combined with the CloudOOP 12000’s made for Hadoop server, can achieves near real time performance for data analytics.

Using the largest capacity Micron SSDs, the CloudOOP 12000 can support up to 96 TBs of storage and is tailored to meet the needs of read-intensive video streaming, latency-sensitive transactional databases and write-intensive logging applications. It also makes an excellent platform for edge computing with the ability to consume large amounts of data very quickly from IOT devices and sensors.

The CloudOOP 12000 is the only server specifically designed for Hadoop, Kafka, Big Data and IOT. It offers 2x the density and up to 35% lower power draw than traditional manufacturers as well as a near 50% increase in data throughput performance. Reducing power draw means a lower data center footprint and significantly reducing your total cost of ownership with an over 90% efficiency rating.

“The CloudOOP1200 is PSSC Lab’s unique platform for enterprise users using applications like Hadoop, Spark, Kafka Streaming. PSSC Labs has already successfully deployed over 100 PBytes for Hadoop using the CloudOOP 12000 platform, and after stringent review of the SSD options on the market, we’ve certified of Micron’s new 5100 series of SSDs, which will allow us to offer our customers a high capacity, high performance, durable system with the absolute lowest cost of ownership, “said Alex Lesser, Vice President of PSSC Labs.

Micron 5100 Series SSD Features include:

High Capacity
Unique range of solutions with up to 8TB of storage in a 2.5-inch form factor and 2TB in an M.2

High Performance
Three models optimized for varying workloads with consistent, steady state random writes at 74,000 IOPS.

Secure Encryption
Built-in AES-256-bit encryption and TCG Enterprise protection with FIPS 140-2 validation – available on the 5100 MAX.

Greater Flexibility
Micron’s FlexPro firmware architecture can be used to actively tune capacity to optimize drive performance and endurance

Best Reliability
Unmatched 99.999%2 quality of service (QoS) compared to spinning media. MTTF of 2 million device hours

CloudOOP 12000 Features Include:

High Processing Power
The CloudOOP 12000 supports up to 2 x Intel Xeon E5 Series processors & up to 256GB high performance memory – get higher performance and reduced computing time

Direct Connect IO technology
Unique design gives each hard drive its own independent path to the motherboard – removing unnecessary components that restrict data pathways and improving data ingestion & IO rates.

Connectivity Options
GigE, 10GigE, 40GigE and Infiniband network connectivity options available. Dual GigE network bandwidth comes standard, with addition network adapters from Intel, Mellanox, Solarflare and others available.

Operating System Compatibility
Supports Microsoft Windows, Red Hat, CentOS, Ubuntu & most other Linux distributions.
All CloudOOP 12000 server configurations service and support from PSSC Lab’s US based, expert in house engineers. Prices for a custom CloudOOP 12000 server start at $5000.

For more information see https://pssclabs.losangeles.dev.buckupstudio.com/products/big-data-servers/.

Source: http://www.prnewswire.com/news-releases/pssc-labs-introduces-affordable-high-performance-all-flash-storage-on-unique-enterprise-server-designed-specifically-for-hadoop-300448604.html

Whitepaper: Loading a Time Series Database at 100 Million Points Per Second.

By marketing@site-a.com

There are many use cases for time series data, and they usually require handling a decent data ingest rate. Rates of more than 10,000 points per second are common and rates of 1 million points per second are not quite as common, but are certainly not unheard of.

Many of these systems are used to monitor critical infrastructure, especially at the higher data rates, where failure of the monitoring systems can lead directly to disastrous failures of the actual system.

It’s important to determine that your time series database will actually work at full loads, so naturally, many organizations will test it at full load. To do this correctly, you have to first fill the database with typically a year or 3 of data, which is where you will suddenly find yourself with a much higher data rate requirement. It makes no sense to test a system with 3 years of data if it takes 3 years to fill the database in preparation for the test.

Thus, if you have a production data rate of 100,000 points per second, you will need to ingest test data at a rate of 100 million points per second if you want to load 3 years of data in less than 2 days. Even a production rate of 10,000 points per second will require several hours to load test data.

Similarly, if you are starting a new database, you probably will want to load significant amounts of historical data, which could be just as stressful as loading the database for testing.

How did we do this?

There are three major components that should be considered:

  1. Loading one point at a time is very slow; batching must be implemented in order to accomplish faster ingestion rates.
  2. The data store needs to be highly performant in order to consistently handle such massive amounts of data.
  3. Well-configured hardware is critical to achieving the highest performance.

First, OpenTSDB was used for this test because it is a well documented time series database and it uses the HBase API for its data store. OpenTSDB also uses the non-relational nature of the HBase API to strong advantage by storing data in a hybrid wide data schema. In our tests, we used one year of data. Sampling one sensor every second for one year generates about 31.5 million points. In OpenTSDB’s storage format, this translates to about 120MB of data per sensor year of data.

Second, MapR-DB was used for this test instead of HBase itself. MapR-DB offers many benefits over HBase, while maintaining the virtues of the HBase API and the idea of data being sorted according to primary key. MapR-DB provides operational benefits such as no compaction delays and automated region splits that do not impact the performance of the database. The tables in MapR-DB can also be isolated to certain machines in a cluster by utilizing the topology feature of MapR. This allowed the testing to easily be scaled to measure the performance on any number of nodes in the cluster. The final differentiator is that MapR-DB is just plain fast, due primarily to the fact that it is tightly integrated into the MapR file system itself, rather than being layered on top of a distributed file system that is layered on top of a conventional file system.

Third, the hardware that was used was fast. The test cluster was designed by PSSC Labs specifically for use with Hadoop systems. The CloudOOP 12000 platform offers a unique, highly optimized system design with a direct data path for disk I/O for up to 12 SATA/SAS devices in a 1u rackmount chassis. This hardware explicitly lacks a raid controller and a backplane, as its intended use is for Hadoop and would otherwise be unused. Here are the specifications for this test setup:

  • Dual Intel E5-2650 v2 Processor (16 total physical cores, 32 hyper-threads)
  • 128GB DDR3 1600
  • 1 – 120GB SSD for the operating system
  • 11 – Western Digital 1TB 7200 RPM 6GB/s
  • Solarflare SFN5152 10GigE network adapter
  • CentOS 6.x
  • Estimated power draw: 240 watt idle / 325 watt 100% load

How were these numbers achieved?

Code was created by modifying portions of OpenTSDB to allow bulk importing of data. This code has been published in the MapR App Gallery and is also available via github at https://github.com/mapr-demos/opentsdb. Fundamentally what is happening is that OpenTSDB stores all points that are for a given hour into a single row in MapR-DB. During normal operations, OpenTSDB inserts each point into the row in a separate column. Once an hour, the entire row is read, columns are collected into a blob and the row is written back to the database. This results in a read and a write per point. In our bulk loading code we directly create the blob for an entire hour of data (3,600 points in our test) and insert that blob into the database in a single put. This single put takes the place of over 7,000 database operations per sensor-hour of data.

We used two edge nodes, both running code to generate and load test data. While these edge nodes were generating data, we observed near 100% CPU utilization on the edge nodes. These edge nodes were not quite hitting their outbound bandwidth limit, as they couldn’t generate data any faster than the CPU bottleneck would allow. To achieve maximum performance, we used eight separate Java virtual machines, each running two threads actively generating the data. Other configurations with more or fewer JVMs and more or fewer threads per JVM gave equal or inferior results.

We limited the output data to only four nodes out of a ten node cluster. Our original plan was to start with four nodes and increase the nodes involved in the test until we achieved our goal of more than 100 million points per second. This test generated random integers that emulated 128 sensors sampled at one second intervals for one year. This yielded over four billion points totaling about 15GB on disk storage. While this data size is less than the memory available on the test nodes, writes to the database are all fully persisted to disk during the test. The ingest rate was approximately 110 million points per second and the data load completed in about 36 seconds. The MapR-DB table was set to the default of 4GB regions, which meant that multiple region splits would occur during the tests. The data replication for the table was set to three, meaning every record was duplicated from the first server to two other servers.

At peak loading, we observed that the two edge nodes together were pushing approximately 1.5GB/s of total data from the edge nodes to the cluster nodes. Inbound data volumes on individual cluster nodes were variable, with peaks near 1.3GB/s. Impressively, the 7200 RPM disks were able to keep up with the network speed and were able to persist data with no issues. Data generation was limited by CPU capacity on the edge nodes. Ingest volume was limited by network bandwidth on the cluster side. It was clear that our loading did not drive the disk subsystems to capacity. Adding more cluster nodes will distribute the network load within the cluster, but we expect that the edge node count is the primary limit on ingestion in the current test configuration.

Could it get any faster?

The storage format of OpenTSDB is reasonable, but could be improved significantly. For instance, time values are stored in the column name for the blob of data. Using a single column name and a compressed binary representation of the data would allow substantial compression of the data. One example of this compression is in the times themselves. With delta coding, the total storage for each time in a blob format could be reduced to less than 1% of the current size. The flags which indicate the time resolution of the timestamps and whether the data is floating point could be repeated only once. For sensor data, resolution is typically limited so that dictionary encodings would likely result in up to 10x compression.

In terms of hardware, the networking bandwidth could be increased, perhaps utilizing dual 10GbE NICs or even higher-end networking gear. The MapR distribution supports transparent application level multiplexing of multiple network interfaces, allowing over 2GB/s of data to be transferred over the network. There is also a possibility that performance would improve with faster disks. These changes would increase the cost of the servers, so it would be worthwhile to understand the cost-benefit of the actual gains.

Performance could also be trivially increased by scaling the system out so that the main data table is spread across more than four nodes.