ENTRIES TAGGED "visualizations"
Cloudera ventures into real-time queries with Impala, data centers are the new landfill, and Jesper Andersen looks at the relationship between art and data.
Here are a few stories from the data space that caught my attention this week.
Cloudera’s Impala takes Hadoop queries into real-time
Cloudera ventured into real-time Hadoop querying this week, opening up its Impala software platform. As Derrick Harris reports at GigaOm, Impala — an SQL query engine — doesn’t rely on MapReduce, making it faster than tools such as Hive. Cloudera estimates its queries run 10 times faster than Hive, and Charles Zedlewski, Cloudera’s cloud VP of products, told Harris that “small queries can run in less than a second.”
Harris notes that Zedlewski pointed out that Impala wasn’t designed to replace business intelligence (BI) tools, and that “Cloudera isn’t interested in selling BI or other analytic applications.” Rather, Impala serves as the execution engine, still relying on software from Cloudera partners — Zedlewski told Harris, “We’re sticking to our knitting as a platform vendor.”
Joab Jackson at PC World reports that “[e]ventually, Impala will be the basis of a Cloudera commercial offering, called the Cloudera Enterprise RTQ (Real-Time Query), though the company has not specified a release date.”
Impala has plenty of competition on this playing field, which Harris also covers, and he notes the significance of all the recent Hadoop innovation:
“I can’t underscore enough how critical all of this innovation is for Hadoop, which in order to add substance to its unparalleled hype needed to become far more useful to far more users. But the sudden shift from Hadoop as a batch-processing engine built on MapReduce into an ad hoc SQL querying engine might leave industry analysts and even Hadoop users scratching their heads.”
You can read more from Harris’ piece here and Jackson’s piece here. Wired also has an interesting piece on Impala, covering the Google F1 database upon which it is based and the Googler Cloudera hired away to help build it.
(Cloudera CEO Mike Olson discussed Impala, Hadoop and the importance of real-time at this week’s Strata Conference + Hadoop World.)
Bitsy Bentley on the work behind a good visualization and why she hopes users will take data interactions for granted.
Because of the size, complexity and density of big data, it’s not always easy to find the important insights hiding in all that information. That’s where data visualization comes into play. A great visualization creates meaning where none existed.
Bitsy Bentley (@bitsybot) is the director of data visualization at GfK Custom Research, where she works with information designers to craft meaningful data experiences for a variety of business audiences. In the following interview, she discusses the space between a “wow” response and an “aha” moment, how her team addresses privacy concerns, and why practice is vital for both visualization creators and viewers.
Bentley will explore related visualization topics during her presentation at Strata Conference + Hadoop World in New York City later this month.
Why are data visualizations an effective way to understand the underlying data?
Bitsy Bentley: There is so much beauty and richness in big datasets, and now that we have enough processing power to harness that richness, it’s little wonder that interest in data visualization is exploding. To quote John Tukey: “The greatest value of a picture is when it forces us to notice what we never expected to see.” My clients find that, whether they’re more concerned with numbers or more concerned with stories, an appropriate visual is integral to their understanding of the data.
Visualization unlocks the serendipity of data analysis. It provides a language that is less intimidating than an overwhelming array of digits. Something as simple as a set of histograms breaking down the distribution of a data store makes it easy to find irregularities and outliers in the data. Read more…
Data artist Jer Thorp on working at the New York Times and the aesthetics of data.
Jer Thorp, data artist in residence at the New York Times, sits at the crossroads of data, art and science. Here he discusses his work at the Times and, more broadly, how aesthetics shape our understanding of data.
LinkedIn's Pete Skomoroch on the key capabilities of data scientists.
In this brief video interview, LinkedIn senior research scientist Pete Skomoroch reveals the three core skills of data scientists.
A deep look at Paul Butler's popular Facebook visualization.
Paul Butler's visualization of Facebook friendships turned a lot of heads recently, and rightfully so. It shows a fascinating connection between virtual relationships and the physical world.
A look at what works in a census visualization.
In this first post in a new series on data visualization, Sébastien Pierre takes a close look at the New York Times' "Mapping America" interactive census map.
Stack Exchange goes in-house, Netflix pays for platforms, survey data gets visualized, and Infochimps acquires Data Marketplace
In this edition of Strata Week: Stack Exchange takes their hardware and software in-house; Neflix explains their adoption of AWS and open source; the New York Times maps out survey and census data; and Infochimps acquires Data Marketplace.
Use Gephi and Python to find your personal communities
Using a bit of Python and the Gephi graph tool, exploring your own Twitter network is a great way to learn about analyzing networks: and the results definitely have a "wow" factor.
Powerful open source graph manipulation
A Photoshop for data, Gephi is a powerful tool for exploring and presenting data as a graph. It's easy to get started with sample data sets, then import your own by generating files in a standard graph format.
Data geekery, visualization and journalism
From deep-diving startup founders to national newspapers, there's a rich vein of wisdom and information in blogs about data. Here's five to get your reading list started.