12 Top Open Source Data Analytics Apps

1. Hadoop

It would be impossible to talk about open source data analytics without mentioning Hadoop. This Apache Foundation project has become nearly synonymous with big data, and it enables large-scale distributed processing of extremely large data sets. A survey conducted by TDWI and SAS found that nearly 60 percent of enterprises expected to have Hadoop clusters in production by the end of 2016.

However, it should be noted that Hadoop on its own doesn‘t enable data analytics. It‘s usually part of a larger solution for gathering insights from big data.

2. Spark

Also an Apache project, Spark promises fast big data processing. In fact, it claims to "run programs up to 100x faster than Hadoop MapReduce in memory, or 10x faster on disk." As a result of this fast performance, it is often used to analyze streaming data or in applications that require interactive analysis capabilities. Companies frequently use it alongside Hadoop or Mesos although it can also run on its own. It has recently experienced a dramatic rise in popularity, and a 2016 survey conducted by Syncsort found that nearly 70 percent of enterprise big data staffers surveyed were interested in Spark.

3. Talend

Unlike the first two projects in this slideshow, Talend is managed by a for-profit company rather than a foundation. As a result, paid support is available. Talend offers a mix of free and paid products. Its free, open source solution is called Talend Open Studio, and it has been downloaded more than 2 million times.

Market research firm Gartner recently named Talend a "Leader" in data integration. The company boasts that it can help enterprises analyze their big data five times faster and at one-fifth the cost compared to competing solutions.

4. Jaspersoft

Like Talend, Jaspersoft comes in multiple editions both free and paid. Its Community edition is free and open source while the Reporting, AWS, Professional and Enterprise editions require a fee but come with support included.

Jaspersoft is an open source business intelligence tool that aims to allow business users to self-serve their own needs. The company claims that its technology powers more than 130,000 apps with embedded BI capabilities.

5. Pentaho

Pentaho describes itself as a "comprehensive data integration and business analytics platform." The company primarily promotes the commercial versions of its software, which are based on the open source Community version. Companies can use it alongside tools like Hadoop and Spark to enable reporting and visualizations for their big data. This software boasts a long list of well-known customers that includes BT, Caterpillar, Nasdaq, The U.S. Dept. of Homeland Security, NOAA, The New York Times, EMC and many others.

6. RapidMiner

RapidMiner claims to be the "#1 open source data science platform," and Gartner named it a leader in its Magic Quadrant report for advanced analytics. It enables self-service predictive analytics and promises lightning-fast performance. Its users include BMW, Lufthansa, Domino‘s Pizza, Sony, Ford, Salesforce, Amnesty International and GE.The complete RadiMiner Platform includes three separate pieces: RapidMiner Studio, RapidMiner Server and RapidMiner Radoop. All three are available under open source or commercial licenses, and commercial prices depend on the number of users.

7. Storm

Used by companies like Yahoo, Twitter, Spotify, Yahoo, Yelp, Flipboard and Groupon, Apache Storm is a real-time big data processing engine. Its website explains, "Storm makes it easy to reliably process unbounded streams of data, doing for real-time processing what Hadoop did for batch processing." Customers can use it with any database and any programming language. It‘s scalable, fault-tolerant and easy to deploy. Users should note however, that Storm has not yet reached the 1.0 release level.

8. H2O

Used by more than 60,000 data scientists at more than 7,000 organizations, H2O claims to be "the world‘s leading open source machine learning platform." Thanks to its in-memory technology, it offers extremely fast performance. It also integrates with many other open source data analytics tools like Hadoop and Spark, and it supports all of the most popular databases. Paid support is available.

In addition to the standard version of H2O, the company also offers Sparkling Water, a version that incorporates Spark, and Steam, and end-to-end artificial intelligence application engine.

9. Lumify

Created by a company called Altamira Technologies, Lumify describes itself as an "open source big data analysis and visualization platform." It makes it easy to create 2D or 3D graphs that show the relationship between entities or to overlay data on maps. For those who are interested in learning more about how it works, the website offers several videos that show Lumify in action, and it also has a demo site that allows users to upload their own data and try out the software.

10. Drill

Apache Drill allows users to use SQL queries for non-relational data storage systems. It supports a range of NoSQL and cloud-based data storage systems, including HBase, MongoDB, MapR-DB, HDFS, MapR-FS, Amazon S3, Azure Blob Storage, Google Cloud Storage and Swift. It also allows users to search through multiple datasets stored with different technologies using a single query. In addition, it supports many popular BI tools.

11. MongoDB

One of the best-known NoSQL databases, MongoDB is an open-source non-relational data storage solution. Its customers include MetLife, the city of Chicago, Expedia, Google, The Weather Channel, BuzzFeed and Facebook. In addition to the free open source version, the company also offers a paid Enterprise version and MongoDB Atlas, a cloud-hosted version. Forrester has named MongoDB a "Leader" for big data NoSQL.

12. SpagoBI

SpagoBI is an open source business intelligence and big data analytics platform. The software is completely free, but paid user support, maintenance, consulting and training are available for purchase. It includes tools for reporting, multidimensional analysis (OLAP), charts, location intelligence, data mining, ETL and more. It also integrates with popular in-memory processing engines and enables real-time processing.

时间: 2024-09-29 18:27:58

12 Top Open Source Data Analytics Apps的相关文章

Toward Scalable Systems for Big Data Analytics: A Technology Tutorial (I - III)

ABSTRACT Recent technological advancement have led to a deluge of data from distinctive domains (e.g., health care and scientific sensors, user-generated data, Internet and financial companies, and supply chain systems) over the past two decades. The

Big Data Analytics for Security(Big Data Analytics for Security Intelligence)

http://www.infoq.com/articles/bigdata-analytics-for-security This article first appeared in the IEEE Security & Privacymagazine and is brought to you by InfoQ & IEEE Computer Society. Enterprises routinely collect terabytes of security-relevant da

CIS 545 - Big Data Analytics

CIS 545 - Big Data Analytics - Fall 2019 Have you ever wondered about (1) what it takes to be a data scientist or "data person", and (2) how sowork?This homework is focused on (1) working with hierarchical data stored in dataframes, (2) traversi

IAB303 Data Analytics Assessment Task

Assessment TaskIAB303 Data Analyticsfor Business InsightSemester I 2019Assessment 2 – Data Analytics NotebookName Assessment 2 – Data Analytics NotebookDue Sun 28 Apr 11:59pmWeight 30% (indicative weighting)Submit Jupyter Notebook via BlackboardRatio

Big Data Analytics and Data Mining 第一天.

今天是上课的第一天.真心很感激导师能让我出来学习.今天突然觉得自己要好好学习英语.并不是上课的时候我看不懂裴教授的课件.而是觉得如果英语不好就很像乡巴佬那样,很难接触到高级的东西. 通过今天的听讲,我感觉对数据挖掘的理解更深刻些. 以前总觉得自己研究生的目标是要好好学习算法,好好学习相关的技术. 现在觉得除了要好好学习算法外,我也期待自己能做出一些研究. 记录下今天讲课的内容. 今天我觉得主要讲了三部分: 1,数据挖掘相关的概念及相关的学术期刊. 从广义上来定义数据挖掘:The art of d

12.2 中的Data Guard Standby 密码文件自动同步 (Doc ID 2307365.1)

Data Guard Standby Automatic Password file Synchronization in 12.2 (Doc ID 2307365.1) APPLIES TO: Oracle Database - Enterprise Edition - Version 12.2.0.1 and laterOracle Database Cloud Schema Service - Version N/A and laterOracle Database Exadata Clo

说一说BDAS(Berkeley Data Analytics Stack)

Strata+Hadoop World 2016在San Jose刚刚结束.对于大数据从业者来讲,这是一定要关注的一个盛会.其中有一个keynote,是Berkeley大学的Michael Franklin的关于BDAS的未来的发展的,非常值得关注,你要问我为什么? BDAS乃是伯克利大学的AMPLab打造的用于大数据的分析的一套开源软件栈,这其中包括了这两年火的爆棚的Spark,也包括了冉冉升起的分布式内存系统Alluxio(Tachyon),当然还包括著名的资源管理的开源软件Mesos.可以

mongodb高可用Replica Set

*************************************************************** 第一部分:系统配置 *************************************************************** ---0.配置yum源 cd /etc/yum.repos.d mv CentOS-Base.repo CentOS-Base.repo.old wget http://mirrors.163.com/.help/CentO

转:python 的开源库

Python在科学计算领域,有两个重要的扩展模块:Numpy和Scipy.其中Numpy是一个用python实现的科学计算包.包括: 一个强大的N维数组对象Array: 比较成熟的(广播)函数库: 用于整合C/C++和Fortran代码的工具包: 实用的线性代数.傅里叶变换和随机数生成函数. SciPy是一个开源的Python算法库和数学工具包,SciPy包含的模块有最优化.线性代数.积分.插值.特殊函数.快速傅里叶变换.信号处理和图像处理.常微分方程求解和其他科学与工程中常用的计算.其功能与软