
Announcement:
Celebrating FTLOD’s 3 year anniversary this month
Covered a diverse range of topics from BBQ and chocolate to alogorithms and graph databases
Future episodes will be much more ad-hoc and when I come across a topic that is interesting
Please stay subscribed
Please reach out on Twitter or LinkedIn to let me know what your favorite episode has been
The Importance of Your Data
Quotes:
With great power comes great responsibility -Amazing Spider Man #15
Data Commercialization
Ford’s CEO recently suggested that the data collected by the company’s financial services arm also represents a valuable, low-overhead asset.1
Not just driving data, but also using data from purchase process such as marital status, income, etc.
However, in desperation to maintain profits, what would some companies do?
Know how your data is being used.
Tim Cook recently criticized Google, FB, and others (not by name) of creating a “‘data industrial complex’ in which our personal information ‘is being weaponized against us with military efficiency.’”2.
Talked about the echo chamber that social networks and algorithms can create
However, this is not all data doomsday
Data is helping us achieve better, deeper, faster insights than ever before
We are bettering our health, optimizing economies, and identifying connections that we never could have before
All this reward comes with some risks that we need to manage and be aware of
Data Breaches
Marriott disclosed a 500MM record breach. Not the biggest, but it hackers had access since 2014.
Names, phone numbers, email addresses, passport numbers, date of birth and arrival and departure information. For millions others, their credit card numbers and card expiration dates were potentially compromised.3
What to do to protect yourself if your data is part of a breach:4
Sign up for services like SpyCloud (it is free)
Change your password – and ideally switch to unique passphrases
Monitor your accounts for suspicious activity
Open a separate credit card for online transactions
Limit the information you share
Avoid saving credit card information on websites
Be vigilant
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
https://threatpost.com/ford-eyes-use-of-customers-personal-data-to-boost-profits/139209/
https://www.nytimes.com/2018/10/26/technology/apple-time-cook-europe.html
https://www.cnn.com/2018/11/30/tech/marriott-hotels-hacked/index.html
https://www.cnn.com/2018/11/30/tech/marriott-breach-what-to-do/index.html
https://answers.kroll.com/
Dec 3, 2018
19 min

In the second part of this two-part episode, we do a data deep dive into a decadent vat of chocolate. We talk about various stats and data with Brian Mikiten, former process engineer and founder of Casa Chocolates in San Antonio, TX. We also cover the types of chocolate and how much of chocolate making is an art vs. a science. See part one for the history of chocolate and an overview of how to make it.
Types
White
Dark
Milk
Ruby – created in 2017 from Ruby cocoa beans by Barry Callebaut in Switzerland
Photo via bakemag.com
Chocolate data
World Chocolate Day is July 7th. US National Chocolate Day is October 28.
The United States accounts for 20% of the world’s chocolate consumption.
On the average Valentine’s Day, nearly $400 million of chocolate is purchased around the world, accounting for 5% of the industry’s total sales.
22% of all chocolate consumed between 8pm and midnight.
Chocolate significantly reduces theta activity in the brain, which is associated with relaxation, which is why we want to eat chocolate when we’re feeling stressed out.
Myth: Chocolate is high in caffeine (contains ~6mg/bar, same as decaf coffee)
More than 70% of Americans prefer milk chocolate
In 2011, Thorntons created the world’s largest chocolate bar, which weighed in at 12,770 lbs. It measured 13 ft. by 13 ft. by 1 ft.
Top companies by sales (via https://www.icco.org)
$ / ton by date
Top 10 World Cocoa Producers
Science / data driven production of chocolate
Equipment used
Variables evaluated / controlled
What is your test process?
Science vs. Art of chocolate making
Bean profiles
Brian’s background and history of Casa Chocolate
What Casa Chocolate’s approach is to making chocolate
Tips for getting started at home
Where people can find out more about Brian and Casa Chocolate
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources
https://www.casachocolates.com
https://en.wikipedia.org/wiki/History_of_chocolate
https://media-cdn.tripadvisor.com/media/photo-s/05/c1/5e/bf/green-acres-chocolate.jpg
https://www.history.com/topics/ancient-americas/history-of-chocolate
http://www.bean-to-bar.co.uk/Category/HowtoMakeyourOwnChocolate/3823 – How to make chocolate
http://dallaschocolate.org/ – Dallas Chocolate Festival
http://chocolatealchemy.com/ – overviews, tips, recipes, equipment info
http://www.chocolatenoise.com/chocolate-today/2017/9/19/my-top-50-bean-to-bar-chocolate-makers-in-the-united-states – List of many bean-to-bar chocolate makers
https://www.amazon.com/SS-Premier-PG-506-Chocolate/dp/B016E1NUZA – Premier Choclate Refiner (Melanger)
https://www.americastestkitchen.com/articles/1000-bean-to-bar-chocolate-trivia-everyone-with-a-sweet-tooth-should-know – Bean-to-Bar overview
https://www.statista.com/chart/3668/the-worlds-biggest-chocolate-consumers/
https://brandongaille.com/26-incredible-chocolate-consumption-statistics/
https://www.icco.org/about-cocoa/chocolate-industry.html
https://www.icco.org/statistics/cocoa-prices/monthly-averages.html?currency=usd&startmonth=01&startyear=2012&endmonth=12&endyear=2018&show=graph&option=com_statistics&view=statistics&Itemid=114&mode=custom&type=1
https://www.candyindustry.com/blogs/14-candy-industry-blog/post/88254-celebrating-country-and-chocolate
https://blog.euromonitor.com/2018/07/global-chocolate-industry.html
https://en.wikipedia.org/wiki/Ruby_chocolate
http://www.sccij.jp/news/overview/detail/article/2018/01/19/nestle-japan-launches-first-ruby-chocolate-product/
https://www.bakemag.com/articles/8656-ruby-chocolate-honored-with-innovation-award-at-sweets-snacks-expo
Oct 31, 2018
31 min

In the first part of this two-part episode, we do a data deep dive into a decadent vat of chocolate. We talk about history and how to make chocolate. In part two, we will talk about various stats and data with Brian Mikiten, former process engineer and founder of Casa Chocolates in San Antonio, TX.
History of chocolate
Evidence dates back as early as 1500 BC
Fermented beverages date back to 350BC
Believed to have originated with Mesoamericans
Made its way to Europe where sugar was added in 16th century
In 1828, Dutch chemist Coenraad Johannes van Houten used alkaline salts to process into “Dutch cocoa”
1847 – J.S. Fry and Sons created the first chocolate bar
1876 – Swiss chocolatier Daniel Peter added milk powder to create milk chocolate
2/3 of cocoa today is produced in Western Africa
Fair trade chocolate certifies that chocolate is not gathered with child or slave labor
Overview of chocolate making
Harvesting – pods contain ~40 cacao beans
Photo via TripAdvisor.com
Roasting
Cracking
Winnowing
Grinding
Conching
Tempering
Molding
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources
https://www.casachocolates.com
https://en.wikipedia.org/wiki/History_of_chocolate
https://media-cdn.tripadvisor.com/media/photo-s/05/c1/5e/bf/green-acres-chocolate.jpg
https://www.history.com/topics/ancient-americas/history-of-chocolate
http://www.bean-to-bar.co.uk/Category/HowtoMakeyourOwnChocolate/3823 – How to make chocolate
http://dallaschocolate.org/ – Dallas Chocolate Festival
http://chocolatealchemy.com/ – overviews, tips, recipes, equipment info
http://www.chocolatenoise.com/chocolate-today/2017/9/19/my-top-50-bean-to-bar-chocolate-makers-in-the-united-states – List of many bean-to-bar chocolate makers
https://www.amazon.com/SS-Premier-PG-506-Chocolate/dp/B016E1NUZA – Premier Choclate Refiner (Melanger)
https://www.americastestkitchen.com/articles/1000-bean-to-bar-chocolate-trivia-everyone-with-a-sweet-tooth-should-know – Bean-to-Bar overview
https://www.statista.com/chart/3668/the-worlds-biggest-chocolate-consumers/
https://brandongaille.com/26-incredible-chocolate-consumption-statistics/
https://www.icco.org/about-cocoa/chocolate-industry.html
https://www.icco.org/statistics/cocoa-prices/monthly-averages.html?currency=usd&startmonth=01&startyear=2012&endmonth=12&endyear=2018&show=graph&option=com_statistics&view=statistics&Itemid=114&mode=custom&type=1
https://www.candyindustry.com/blogs/14-candy-industry-blog/post/88254-celebrating-country-and-chocolate
https://blog.euromonitor.com/2018/07/global-chocolate-industry.html
https://en.wikipedia.org/wiki/Ruby_chocolate
http://www.sccij.jp/news/overview/detail/article/2018/01/19/nestle-japan-launches-first-ruby-chocolate-product/
https://www.bakemag.com/articles/8656-ruby-chocolate-honored-with-innovation-award-at-sweets-snacks-expo
Oct 24, 2018
1 hr 3 min

Intro
Greg’s Background
Intro to DevOps
Tools you’ve used
Intro to the report & this year vs. previous years
Feels more general and high-level than some of the previous reports (no MTTR mentioned for instance)
Who took the survey?
Surveyed over 30,000 people in 7 years (~4,300 / yr)
Technology is overrepresented – 38% of total respondents
Energy & Resources was only 2%
Tech + FS = 50%
Infosec = only 3% of people
29% were dedicated DevOps (14% IT, 15% Dev/Eng)
Keywords
Data is mentioned 43 times in the report
Security = 65x
Agile = 7
DevOps = 328
Top 10 words in Word Cloud: (removed Puppet | State of DevOps footer on each page)
245DevOps
210teams
149Stage
123practices
111can
99team
79organizations
75services
73business
69success
No DataOps, no SecOps or DevSecOps
C-suite seems out of touch with conditions on the ground
Differences in perception – p. 30
Sometimes overstate team’s opinion by a factor of 2x
Stages in the report:
First
Stage 0: Build the foundation
Second
Stage 1: Normalize the technology stack
Stage 2: Standardize and reduce variability
Stage 3: Expand DevOps practices
Third
Stage 4: Automate infrastructure delivery
Stage 5: Provide self-service capabilities
CAMS = Culture, Automation, Measurability, Sharing
Principal Industries:
Top: Tech, Financial Services, Manufacturing/Industry
Bottom: Non-Profit, Energy/Resources, Media
Trend: Most to least competition?
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
https://puppet.com/resources/whitepaper/state-of-devops-report
https://devops-research.com/ – DORA – DevOps Research & Assessment, authors of the report
https://www.wordclouds.com/
https://www.linkedin.com/in/gregwwalters/
Sep 30, 2018
45 min

Learn about Cursor, a new platform for collaboration around data, hosted platforms and BI artifacts. I sat down with Adam Weinstein, CEO and Co-Founder of Cursor, to learn about the platform.
About Cursor
Cursor offers a data search and analytics hub that makes disparate data accessible and actionable, enabling technical and business users alike to effortlessly get answers, collaborate and gain insights. Founded by a trio of data leaders from Salesforce, LinkedIn, and Pandora, Cursor’s easy-to-deploy software has been adopted by teams at Apple, Atlassian, Deloitte, Incedo, LinkedIn, NovumRx, and Slack. Cursor is based in San Francisco, CA.
Cursor Press
https://techcrunch.com/2018/05/30/cursor-looks-to-build-a-search-tool-for-any-internal-database-with-2m-in-new-funding/
https://www.businessinsider.com/adam-weinstein-linkedin-cursor-2018-7
Topics:
What is Adam’s background?
How BI has evolved over the past 10-20 years.
What some of the most pressing challenges are for organizations today?
What should people being doing today, outside of a specific tool, to get better at collaborating?
How can Cursor help with those challenges?
How is content secured on the platform? (separating data from metadata)
Where can people find out more about Cursor?
What’s next for Cursor as far as features or a roadmap?
What are some tools that Adam can’t live without in your daily work?
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Aug 30, 2018
40 min

July 2018 News Roundup
This month’s episode is a roundup of news from a variety of sources covering three main topics:
BI / Dataviz Tools
Databases and Platforms
Tools and Frameworks
Note: Most of the text extracts below are direct quotations from new sources cited in the source list at the bottom of these show notes. This episode is a compilation from those sources.
BI / Dataviz Tools
PowerBI enhancements (7/12/18)
Microsoft has updated its Power BI analytics service in an effort to expand data prep capabilities and unify data analytics across platforms.
“Using the Power Query experience familiar to millions of Power BI Desktop and Excel users, business analysts can ingest, transform, integrate and enrich big data directly in the Power BI web service – including data from a large and growing set of supported on-premises and cloud-based data sources, such as Dynamics 365, Salesforce, Azure SQL Data Warehouse, Excel and SharePoint,” the post reads.
Power BI now supports data in Azure Data Lake Storage, and integrates with SQL Server Analysis Services and SQL Server Reporting Services.
Microsoft today announced the general availability of Visio Visual for Power BI. Based on the feedback collected from the customers during the preview period, Microsoft has made the following changes to the Visio Visual:
Support for Power BI Mobile app
The ability to change the diagram link embedded earlier and to copy an embedded link to the clipboard
Configurable auto-zoom settings that can be turned on and off
Support for complex diagrams using layers
Overall performance improvements
Tableau acquires Empirical Systems
Tableau last month announced the acquisition of Empirical Systems, an artificial intelligence (AI) startup with an automated discovery and analysis engine designed to spot influencers, key drivers, and exceptions in data.
Looker Enhances Data Science Capability with Integration for Google Cloud BigQuery ML
With Looker and BQML, data teams can now save time and eliminate unnecessary processes by creating machine learning (ML) models directly in Google BigQuery via Looker – without the need to transfer data into additional ML tools. BQML predictive functionality will also be integrated into new or existing Looker Blocks allowing users to surface predictive measures in dashboards and applications.
DBs and Platforms
MemSQL Unveil Significant Update to Database for Real-time Modern Applications and Analytical Systems (Version 6.5 released)
Queries are now up to four times faster than the previous MemSQL version (which was already 10x faster than legacy database providers), enabling insights in milliseconds across billions of rows.
New automated workload optimization capabilities provide a consistent database response under ultra-high concurrency without the need for manual tuning or specialized DBA resources.
Additions to the MemSQL industry-leading “transform-as-you-ingest” capabilities allow customers to use stored procedures for in-database transformations to easily build real-time data pipelines.
Resource optimization improvements for multi-tenant deployments deliver greater control and scalability for varied database sizes whether on-premises or in the cloud.
Hortonworks Data Platform 3.0
Even a Hadoop stalwart such as Hortonworks Inc. sees the writing on the wall, which is why, in its recent 3.0 release, it emphasized heterogeneous object storage. The new Hortonworks Data Platform 3.0 supports data storage in all of the major public-cloud object stores, including Amazon S3, Azure Storage Blob, Azure Data Lake, Google Cloud Storage and AWS Elastic MapReduce File System.
HDP’s latest storage enhancements include a consistency layer, NameNode enhancements to support scale-out persistence of billions of files with lower storage overhead, and storage-efficiency enhancements such as support for erasure coding across heterogeneous volumes. HDP workloads access non-HDFS cloud storage environments via the Hadoop Compatible File System API.
My thoughts: Are Hadoop and HDFS Dying???
As we are heading into the fourth industrial revolution, HDP 3.0 is a giant leap for the Big Data ecosystem, with major changes across the stack and expanded eco-system (Deep Learning and 3rd Party Dockerized Apps). HDP 3.0 can be deployed both on-premise and in the major cloud platforms – AWS, Microsoft Azure, and Google Cloud. Many of the HDP 3.0 new features are based on Apache Hadoop 3.1 and include containerization, GPU support, Erasure Coding and Namenode Federation. In order to provide a Trusted Data Lake, we are installing Apache Ranger and Apache Atlas by default with HDP 3.0. In order to streamline the stack, we have removed components such as Apache Falcon, Apache Mahout, Apache Flume, and Apache Hue, and absorbed Apache Slider functionalities into Apache YARN.
Tools and Frameworks
Python 3.7.0 is now available
Data classes that reduce boilerplate when working with data in classes.
A potentially backward-incompatible change involving the handling of exceptions in generators.
A “development mode” for the interpreter.
Nanosecond-resolution time objects.
UTF-8 mode that uses UTF-8 encoding by default in the environment.
A new built-in for triggering the debugger.
Easier access to debuggers through a new breakpoint() built-in
Simple class creation using data classes
Customized access to module attributes
Improved support for type hinting
Higher precision timing functions
More importantly, Python 3.7 is fast.
Each new release of Python comes with a set of optimizations. In Python 3.7, there are some significant speed-ups, including:
There is less overhead in calling many methods in the standard library.
Method calls are up to 20% faster in general.
The startup time of Python itself is reduced by 10-30%.
Importing typing is 7 times faster.
You can easily get an idea of how much time the imports in your script takes, using -X importtime:
Apache OpenNLP 1.9.0 released
The Apache OpenNLP team is pleased to announce the release of Apache OpenNLP 1.9.0.
The Apache OpenNLP library is a machine learning based toolkit for the processing of natural language text.
It supports the most common NLP tasks, such as tokenization, sentence segmentation, part-of-speech tagging, named entity extraction, chunking, parsing, and coreference resolution.
Apache OpenNLP 1.9.0 binary and source distributions are available for download from our download page: download page
The OpenNLP library is distributed by Maven Central as well. See the Maven Dependency page for more details: Maven Dependency
What’s new in Apache OpenNLP 1.9.0
This release introduces new features, improvements and bug fixes. Java 1.8 and Maven 3.3.9 are required.
Additionally the release contains the following changes:
Brat Document Parser should support name type filters
Brat format support fails on multi fragment annotations
Remove MD5 hashes from Release process
Use String[] instead of StringList in LanguageModel API
BRAT Annotator service Fails to start
Token model creation fails without at least one <SPLIT> tag
Update Penn Treebank URL
Explain the new format of feature generator XML config
Unify code to sum up input context features
FeatureGeneratorUtil can recognize Japanese Hiragana and Katakana letters
TensorFlow 1.9.0
Updated docs for tf.keras: New Keras-based get started and programmers guide page.
Update tf.keras to the Keras 2.1.6 API.
Added tf.keras.layers.CuDNNGRU and tf.keras.layers.CuDNNLSTM layers. Try it.
Adding support of core feature columns and losses to gradient boosted trees estimators.
The python interface for the TFLite Optimizing Converter has been expanded, and the command line interface (AKA: toco, tflite_convert) is once again included in the standard pip installation.
Improved data-loading and text processing with:
tf.decode_compressed
tf.string_strip
tf.strings.regex_full_match
Added experimental support for new pre-made Estimators:
tf.contrib.estimator.BaselineEstimator
tf.contrib.estimator.RNNClassifier
tf.contrib.estimator.RNNEstimator
The distributions.Bijector API supports broadcasting for Bijectors with new API changes.
PYPL Language Rankings: Python ranks #1, R at #7 in popularity
The new PYPL Popularity of Programming Languages (June 2018) index ranks Python at #1 and R at #7.
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
https://www.techrepublic.com/article/microsoft-power-bi-expansion-aims-to-help-analysts-leverage-business-data-more-easily/
https://mspoweruser.com/visio-visual-for-power-bi-now-available-for-everyone/
https://appsource.microsoft.com/en-us/product/office/WA104381132?src=office&corrid=5ebc6738-bf41-4363-bfc9-d46077451402&omexanonuid=19a90565-f1ed-4800-a00b-8ed084f0fe61
https://www.zdnet.com/article/tableau-takes-next-steps-toward-smart-analytics/
https://insidebigdata.com/2018/07/29/memsql-unveil-significant-update-database-real-time-modern-applications-analytical-systems/
http://blog.revolutionanalytics.com/2018/07/ai-roundup-july-2018.html
https://pythoninsider.blogspot.com/2018/06/python-3.html
https://www.python.org/downloads/release/python-370/
https://www.infoworld.com/article/3252852/python/whats-new-in-python-37.html
https://realpython.com/python37-new-features/
http://blog.revolutionanalytics.com/2018/06/pypl-programming-language-trends.html
http://pypl.github.io/PYPL.html
https://insidebigdata.com/2018/07/29/looker-enhances-data-science-capability-integration-google-cloud-bigquery-ml/
https://github.com/tensorflow/tensorflow/releases/tag/v1.9.0
https://opennlp.apache.org/news/release-190.html
https://siliconangle.com/2018/07/09/hadoops-star-dims-era-cloud-object-data-storage-stream-computing/
https://hortonworks.com/blog/announcing-general-availability-hortonworks-data-platform-3-0-0-ambari-2-7-0-smartsense-1-5-0/
Jul 31, 2018
15 min

Is Data the New Oil?
Concept originated by Clive Humby, the British mathematician who established Tesco’s Clubcard loyalty program. Humby highlighted the fact that, although inherently valuable, data needs processing, just as oil needs refining before its true value can be unlocked.
Why it is the new oil
Valuable commodity
Different uses among many applications
Currently the big buzz of most large companies (Google, Facebook, Apple, etc.)
Quantity is generally better in both
AI is the darling of so many industries right now, and it is entirely dependent on data
There are ethical concerns with how we source and use this, just like there were and are geopolitical and ethical concerns with how we source and use Oil
Certain things cannot function (currently) without oil (passenger airplanes, boats)
Same with data: Oil & Gas, Netflix, Agriculture, Manufacturing, Healthcare, a general enabler
Why it isn’t the new oil
Oil is finite, but data is not
Rob’s Counterpoint: There is a shelf life on data that makes it less usable over time
Data does not have a standard price benchmark like oil
Not a physical asset; can be duplicated or shared relatively easily
Oil requires huge amounts of resources to recover and transport
Rob’s Counterpoint: building a successful “app” with the scale to generate meaningful data does have some costs, albeit not the scale of oil
Data is more useful the more that it is used, whereas oil loses energy the more it is used/processed
Rob’s Counterpoint: Oil is not useful by itself to most people; it’s really the product oil becomes or enables that is useful
The Data of Oil
Difference between operating on surface vs. subsea: small tubing error occurs…
Surface: 2-3 hours downtime; a few thousand $$ to fix
Subsea: 3 months downtime, $40-50mm to fix, not including lost revenue due to deferred production (ex. 15,000 bpd well * $67/barrel * 90 days = $90.45mm)
A good sized offshore platform generates revenue greater than the entire country of Belize ($2.3bn vs. $1.8bn)
Of all the oil we can find, we generally only recover 10-20% in a field with current technology
45-50% of oil generated in the US is used for transportation
US consumption per day is about 2 ½ gallons of crude oil / day / person
The U.S. has 4% of the world’s population but uses 25% of the world’s oil
Total daily oil consumption around the world is 84,249,000 barrels/day
Top 3 countries by proven oil reserves are: Venezuela, Saudi Arabia, Canada; US is #10
Gas is 12,200 Wh/kg vs. Li-Ion at 265 Wh/kg (~46x more energy dense)
MTTF (Mean time to failure) – 500 years on some parts – needed to operate in subsea environments for 30 years
Area of dinner plate = 10.5”
Area = Pi * R^2 * 20,000 PSI
Area = ¼ * Pi * D^2 * 20,000 PSI
0.25 * pi * 10.5 * 10.5 * 20,000 = 1.73180295029137E6
= 1,731,802 pounds on a single dinner plate (equivalent to ~9 737 Jets)
Length Records
Analogy: Standing on top of the Empire State Building in NYC and trying to put a straw in a coke can sitting on the sidewalk below
Deepest Well (scientific study) = Kola Superdeep Borehole= 40,230 ft.
CHAYVO WELL – SAKHALIN-I PROJECT-The current world record holder for longest well; depth of 44,291 feet with a horizontal reach of 39,478 feet
DEEPWATER HORIZON – drilled the deepest oil well in history. The well was drilled to 35,050 vertical depth
Well depth by year
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
Is Data the New Oil?
https://www.economist.com/news/leaders/21721656-data-economy-demands-new-approach-antitrust-rules-worlds-most-valuable-resource
https://www.forbes.com/sites/bernardmarr/2018/03/05/heres-why-data-is-not-the-new-oil/#326090363aa9
https://www.forbes.com/sites/howardbaldwin/2014/03/28/big-datas-big-impact-across-industries/#3ebd563856b9
https://hbr.org/2018/01/is-your-companys-data-actually-valuable-in-the-ai-era
https://ischool.syr.edu/infospace/2018/02/26/why-data-is-the-new-oil/
The Data of Oil Sources:
https://www.bp.com/content/dam/bp/en/corporate/pdf/energy-economics/energy-outlook/bp-energy-outlook-2018.pdf
https://oges.info/library/152237/World_s-Deepest-Drilled-Wells
http://www.investinganswers.com/investment-ideas/commodities-precious-metals/50-surprising-facts-you-never-knew-about-oil-2692
https://www.dosomething.org/facts/11-facts-about-oil
https://oilprice.com/Energy/Crude-Oil/How-Much-Crude-Oil-Do-You-Consume-On-A-Daily-Basis.html
https://en.wikipedia.org/wiki/List_of_countries_by_proven_oil_reserves
https://xtronics.com/wiki/Energy_density.html
https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal)
Other Data is the New Oil Sources:
https://techcrunch.com/2018/03/27/data-is-not-the-new-oil/amp/
http://ana.blogs.com/maestros/2006/11/data_is_the_new.html
https://www.weforum.org/agenda/2018/01/data-is-not-the-new-oil/
https://www.wired.com/insights/2014/07/data-new-oil-digital-economy/
https://www.domo.com/learn/data-never-sleeps-5?aid=DPR072517
https://mobile.cioinsight.com/it-strategy/big-data/data-is-the-new-oil.html
https://cloudtweaks.com/2017/01/data-revolution-new-oil/
http://www.oedigital.com/component/k2/item/11383-data-as-the-new-oil
http://amp.timeinc.net/fortune/2016/07/11/data-oil-brainstorm-tech
http://www.freepressjournal.in/editorspick/ajit-ranade-economic-conclave-data-is-the-new-oil/1108742
Other Oil Sources:
https://seekingalpha.com/article/4136900-2018-oil-production-forecast-explained
http://www.rrc.state.tx.us/oil-gas/research-and-statistics/
https://www.eia.gov/petroleum/drilling/
Jun 29, 2018
59 min

Today we’re back with another guest from the Netherlands. I’m not sure what it is about the Dutch, but they’ve been on a roll with some helpful thought leadership when it comes to data. My guest is Rick van der Lans, a highly-respected analyst, consultant, author, and international lecturer specializing in data warehousing, business intelligence, big data, and database technology.
I came across one of Rick’s whitepapers a few months ago on data virtualization. We got in touch and sat down to talk more in depth about the topic. Rick has a lot of data street cred. For many years, he has served as the chairman of the annual European Enterprise Data and Business Intelligence Conference in London and the annual Data Warehousing and Business Intelligence Summit in The Netherlands. He has written tons of articles, blogs, and several books, including the first book on SQL. There will be links to some of the places and things Rick has written and other info in the show notes below.
Topics:
Rick’s background: author, blogger, consultant – worked on data virtualization (DV) for last 6-7 years
How did Rick get interested in DV?
Classical data warehouses vs. logical data warehouses
What is bi-modal BI? (term introduced by Gartner in 2014)
Agile/Self-Service vs. longer, more cautious approach
Bi-modal BI vs. the Data Quadrant
Comparison of major Data Virtualization Vendors
Denodo
Tibco DV Manager (bought from Cisco recently)
Red Hat
Data Virtuality Ultrarep
Others (AtScale, Cero, StoneBond, IBM – new entry acquired from Rocket Software)
Some are more mature, some are newer (Denodo vs. Tibco = green apples vs. red apples)
Companies rolling their own DV (in-memory / views vs. a dedicated tool)
DV products are not DB views on steroids
Lineage / impact analysis and other features
Caching vs. materialization – can store cached data in a virtual table in an intermediary data store. Can be help performance or prevent interference from a transactional source (keeping results consistent for an entire week).
How DV can help organizations that are struggling
How DV may not be a silver bullet
How are different industries embracing these principles?
What patterns do you see in companies embracing these principles?
What companies should not use this? DV not great at this time on unstructured audio / video, auto-tagging of images
Why a classical DWH experienced person may fail at DV
What are the warning signs that a DV is going off the rails?
Fuzzy logic needed to combine disparate sources
Not an integration cureall
How you deploy these with projects
How to get started? (pick a single, sexy report as a starting point)
Where do you go next? (how to unify other data delivery systems, data marketplaces, API gateways)
How to avoid misconceptions about DV (it is slow, only about integration, etc.)
How to contact Rick
The first book on SQL
Places to find Rick’s work:
He has published blogs for the following websites:
TechTarget.com
B-eye-Network
Business-analytics.biz
He has written the following books:
Data Virtualization for Business Intelligence Systems
Introduction to SQL (has sold over 100,000 copies and has been translated in many languages)
SQL for MySQL Developers
The SQL Guide to SQLite
The SQL Guide to Ingres
The SQL Guide to Pervasive PSQL
Music:
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
https://www.denodo.com/en/system/files/document-attachments/rick_van_der_lans_bimodal_ldw.pdf
https://solutionsreview.com/data-integration/4-data-virtualization-vendors-to-watch-in-2017/
https://reprints.forrester.com/#/assets/2/545/’RES133042’/reports (Forrester Wave for Data Virtualization)
http://www.datavirtualizationblog.com/achieving-lightning-fast-performance-logical-data-warehouse/
http://www.r20.nl/rick_f__van_der_lans1.htm
May 27, 2018
48 min

My guest in today’s episode is Ronald Damhof (@ronalddamhof), the creator of the Data Quadrant. This quadrant is a sense-making framework in the complex word of data that enables a common frame of reference between managers, domain experts and engineers. This model is used by many organisations to formulate data strategy and justify investments in the data domain. It is used as the strategic underpinning for a data architecture, it guides the ‘rules of the game’ and it separates the fundamental concerns in data. Furthermore, it explains how an organisation can toggle the need to innovate with data and the need to deploy and use data at scale, repeatedly, safe, lawful, with constant quality and robust.
Topics:
Ronald background as a “data fundamentalist”
His concept of a full scale data architect
The push / pull point, from 1950s Toyota, applied to data
Development styles from systemic to opportunistic
Data Vault’s influence on the quadrant
Where data modelling (Q1) and data lakes (Q3) fit into the quadrants
Where should you start? Q1/Q2 or Q3/Q4
90% of organizations in the Netherlands are using Data Vault
General recommendations on tools by quadrant:
Q1
Automation – Wherescape or custom
Federalization – mainly still RDBMS
Q2
API’ing the data
Losing faith in datasets and data marts
Q3
Fast infra
Doesn’t believe Hadoop is a good fit for most orgs.
Likes fast analytical DBs like Vertica or MonetDB?
Q4
Open source
R, Python, Git, Dataiku
Abstraction layer away from code is helpful
Azure Platform
Ronald Damhof’s background:
Primary degree in Economics
Certified Data Vault Grand Master
Data Architect at the Dutch Central Bank in the Netherlands
Music
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
https://twitter.com/ronalddamhof
http://prudenza.typepad.com/ (Ronald’s Blog)
https://www.linkedin.com/in/ronalddamhof/ (Ronald’s LinkedIn Profile)
http://www.tdwi-konferenz.de/tdwi2018/startseite-englisch/program/conference-program/sessiondetails/action/detail/session/di-31-3/title/it-is-all-about-the-data-a-managerial-perspective.html
http://www.b-eye-network.com/blogs/damhof/archives/2013/08/4_quadrant_mode.php
http://prudenza.typepad.com/dwh/2015/06/make-data-management-a-live-issue-for-discussion-throughout-the-organization.html
https://www.linkedin.com/in/thefullscaledataarchitect/
Apr 28, 2018
55 min

Introductory Product Models:
Only partner implementations (BluePrism)
Limited Features (WorkFusion)
Customer Revenue Limited (UI Path)
Single License (Softomotive)
Music
Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Sources:
https://irpaai.com/definition-and-benefits/
https://www.edgeverve.com/wp-content/uploads/2017/02/forrester-wave-robotic-process-automation.pdf
https://www.gartner.com/doc/reprints?id=1-3U26FK2&ct=170222&st=sb
http://www.uipath.com/hubfs/News_photos/Forrester_Wave_RPA_Report.png?t=1522186102828
http://images.abbyy.com/India/market_guide_for_robotic_pro_319864%20(002).pdf
https://www.uipath.com/community
https://www.workfusion.com/rpaexpress
https://idm.net.au/article/0011800-which-rpa-software-should-i-use
Mar 30, 2018
27 min
Load more
