For the Love of Data
For the Love of Data
For the Love of Data
We love data and how it intersects with news, products, technologies, and companies. Listen to our podcast and join the discussion to stay informed on the latest and greatest in the world of BI and analytics.
E035 – Your Data and an Announcement
Announcement: Celebrating FTLOD’s 3 year anniversary this month Covered a diverse range of topics from BBQ and chocolate to alogorithms and graph databases Future episodes will be much more ad-hoc and when I come across a topic that is interesting Please stay subscribed Please reach out on Twitter or LinkedIn to let me know what your favorite episode has been   The Importance of Your Data Quotes: With great power comes great responsibility -Amazing Spider Man #15 Data Commercialization Ford’s CEO recently suggested that the data collected by the company’s financial services arm also represents a valuable, low-overhead asset.1 Not just driving data, but also using data from purchase process such as marital status, income, etc. However, in desperation to maintain profits, what would some companies do? Know how your data is being used. Tim Cook recently criticized Google, FB, and others (not by name) of creating a “‘data industrial complex’ in which our personal information ‘is being weaponized against us with military efficiency.’”2. Talked about the echo chamber that social networks and algorithms can create However, this is not all data doomsday Data is helping us achieve better, deeper, faster insights than ever before We are bettering our health, optimizing economies, and identifying connections that we never could have before All this reward comes with some risks that we need to manage and be aware of Data Breaches Marriott disclosed a 500MM record breach. Not the biggest, but it hackers had access since 2014. Names, phone numbers, email addresses, passport numbers, date of birth and arrival and departure information. For millions others, their credit card numbers and card expiration dates were potentially compromised.3 What to do to protect yourself if your data is part of a breach:4 Sign up for services like SpyCloud (it is free) Change your password – and ideally switch to unique passphrases Monitor your accounts for suspicious activity Open a separate credit card for online transactions Limit the information you share Avoid saving credit card information on websites Be vigilant Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: https://threatpost.com/ford-eyes-use-of-customers-personal-data-to-boost-profits/139209/ https://www.nytimes.com/2018/10/26/technology/apple-time-cook-europe.html https://www.cnn.com/2018/11/30/tech/marriott-hotels-hacked/index.html https://www.cnn.com/2018/11/30/tech/marriott-breach-what-to-do/index.html https://answers.kroll.com/  
Dec 3, 2018
19 min
E034 – Using Data to Make Perfect Chocolate – Part 2
In the second part of this two-part episode, we do a data deep dive into a decadent vat of chocolate. We talk about various stats and data with Brian Mikiten, former process engineer and founder of Casa Chocolates in San Antonio, TX. We also cover the types of chocolate and how much of chocolate making is an art vs. a science. See part one for the history of chocolate and an overview of how to make it. Types White Dark Milk Ruby – created in 2017 from Ruby cocoa beans by Barry Callebaut in Switzerland Photo via bakemag.com Chocolate data World Chocolate Day is July 7th. US National Chocolate Day is October 28. The United States accounts for 20% of the world’s chocolate consumption. On the average Valentine’s Day, nearly $400 million of chocolate is purchased around the world, accounting for 5% of the industry’s total sales. 22% of all chocolate consumed between 8pm and midnight. Chocolate significantly reduces theta activity in the brain, which is associated with relaxation, which is why we want to eat chocolate when we’re feeling stressed out. Myth: Chocolate is high in caffeine (contains ~6mg/bar, same as decaf coffee) More than 70% of Americans prefer milk chocolate In 2011, Thorntons created the world’s largest chocolate bar, which weighed in at 12,770 lbs. It measured 13 ft. by 13 ft. by 1 ft. Top companies by sales (via https://www.icco.org) $ / ton by date Top 10 World Cocoa Producers Science / data driven production of chocolate Equipment used Variables evaluated / controlled What is your test process? Science vs. Art of chocolate making Bean profiles Brian’s background and history of Casa Chocolate What Casa Chocolate’s approach is to making chocolate Tips for getting started at home Where people can find out more about Brian and Casa Chocolate Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources https://www.casachocolates.com https://en.wikipedia.org/wiki/History_of_chocolate https://media-cdn.tripadvisor.com/media/photo-s/05/c1/5e/bf/green-acres-chocolate.jpg https://www.history.com/topics/ancient-americas/history-of-chocolate http://www.bean-to-bar.co.uk/Category/HowtoMakeyourOwnChocolate/3823 – How to make chocolate http://dallaschocolate.org/ – Dallas Chocolate Festival http://chocolatealchemy.com/ – overviews, tips, recipes, equipment info http://www.chocolatenoise.com/chocolate-today/2017/9/19/my-top-50-bean-to-bar-chocolate-makers-in-the-united-states – List of many bean-to-bar chocolate makers https://www.amazon.com/SS-Premier-PG-506-Chocolate/dp/B016E1NUZA – Premier Choclate Refiner (Melanger) https://www.americastestkitchen.com/articles/1000-bean-to-bar-chocolate-trivia-everyone-with-a-sweet-tooth-should-know – Bean-to-Bar overview https://www.statista.com/chart/3668/the-worlds-biggest-chocolate-consumers/ https://brandongaille.com/26-incredible-chocolate-consumption-statistics/ https://www.icco.org/about-cocoa/chocolate-industry.html https://www.icco.org/statistics/cocoa-prices/monthly-averages.html?currency=usd&startmonth=01&startyear=2012&endmonth=12&endyear=2018&show=graph&option=com_statistics&view=statistics&Itemid=114&mode=custom&type=1 https://www.candyindustry.com/blogs/14-candy-industry-blog/post/88254-celebrating-country-and-chocolate https://blog.euromonitor.com/2018/07/global-chocolate-industry.html https://en.wikipedia.org/wiki/Ruby_chocolate http://www.sccij.jp/news/overview/detail/article/2018/01/19/nestle-japan-launches-first-ruby-chocolate-product/ https://www.bakemag.com/articles/8656-ruby-chocolate-honored-with-innovation-award-at-sweets-snacks-expo
Oct 31, 2018
31 min
E033 – Using Data to Make Perfect Chocolate – Part 1
In the first part of this two-part episode, we do a data deep dive into a decadent vat of chocolate. We talk about history and how to make chocolate. In part two, we will talk about various stats and data with Brian Mikiten, former process engineer and founder of Casa Chocolates in San Antonio, TX. History of chocolate Evidence dates back as early as 1500 BC Fermented beverages date back to 350BC Believed to have originated with Mesoamericans Made its way to Europe where sugar was added in 16th century In 1828, Dutch chemist Coenraad Johannes van Houten used alkaline salts to process into “Dutch cocoa” 1847 – J.S. Fry and Sons created the first chocolate bar 1876 – Swiss chocolatier Daniel Peter added milk powder to create milk chocolate 2/3 of cocoa today is produced in Western Africa Fair trade chocolate certifies that chocolate is not gathered with child or slave labor Overview of chocolate making Harvesting – pods contain ~40 cacao beans Photo via TripAdvisor.com Roasting Cracking Winnowing Grinding Conching Tempering Molding Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources https://www.casachocolates.com https://en.wikipedia.org/wiki/History_of_chocolate https://media-cdn.tripadvisor.com/media/photo-s/05/c1/5e/bf/green-acres-chocolate.jpg https://www.history.com/topics/ancient-americas/history-of-chocolate http://www.bean-to-bar.co.uk/Category/HowtoMakeyourOwnChocolate/3823 – How to make chocolate http://dallaschocolate.org/ – Dallas Chocolate Festival http://chocolatealchemy.com/ – overviews, tips, recipes, equipment info http://www.chocolatenoise.com/chocolate-today/2017/9/19/my-top-50-bean-to-bar-chocolate-makers-in-the-united-states – List of many bean-to-bar chocolate makers https://www.amazon.com/SS-Premier-PG-506-Chocolate/dp/B016E1NUZA – Premier Choclate Refiner (Melanger) https://www.americastestkitchen.com/articles/1000-bean-to-bar-chocolate-trivia-everyone-with-a-sweet-tooth-should-know – Bean-to-Bar overview https://www.statista.com/chart/3668/the-worlds-biggest-chocolate-consumers/ https://brandongaille.com/26-incredible-chocolate-consumption-statistics/ https://www.icco.org/about-cocoa/chocolate-industry.html https://www.icco.org/statistics/cocoa-prices/monthly-averages.html?currency=usd&startmonth=01&startyear=2012&endmonth=12&endyear=2018&show=graph&option=com_statistics&view=statistics&Itemid=114&mode=custom&type=1 https://www.candyindustry.com/blogs/14-candy-industry-blog/post/88254-celebrating-country-and-chocolate https://blog.euromonitor.com/2018/07/global-chocolate-industry.html https://en.wikipedia.org/wiki/Ruby_chocolate http://www.sccij.jp/news/overview/detail/article/2018/01/19/nestle-japan-launches-first-ruby-chocolate-product/ https://www.bakemag.com/articles/8656-ruby-chocolate-honored-with-innovation-award-at-sweets-snacks-expo
Oct 24, 2018
1 hr 3 min
E032 – 2018 State of DevOps Report
Intro Greg’s Background Intro to DevOps Tools you’ve used Intro to the report & this year vs. previous years Feels more general and high-level than some of the previous reports (no MTTR mentioned for instance) Who took the survey? Surveyed over 30,000 people in 7 years (~4,300 / yr) Technology is overrepresented – 38% of total respondents Energy & Resources was only 2% Tech + FS = 50% Infosec = only 3% of people 29% were dedicated DevOps (14% IT, 15% Dev/Eng) Keywords Data is mentioned 43 times in the report Security = 65x Agile = 7 DevOps = 328 Top 10 words in Word Cloud: (removed Puppet | State of DevOps footer on each page) 245DevOps 210teams 149Stage 123practices 111can 99team 79organizations 75services 73business 69success No DataOps, no SecOps or DevSecOps C-suite seems out of touch with conditions on the ground Differences in perception – p. 30 Sometimes overstate team’s opinion by a factor of 2x Stages in the report: First Stage 0: Build the foundation Second Stage 1: Normalize the technology stack Stage 2: Standardize and reduce variability Stage 3: Expand DevOps practices Third Stage 4: Automate infrastructure delivery Stage 5: Provide self-service capabilities CAMS = Culture, Automation, Measurability, Sharing Principal Industries: Top: Tech, Financial Services, Manufacturing/Industry Bottom: Non-Profit, Energy/Resources, Media Trend: Most to least competition? Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: https://puppet.com/resources/whitepaper/state-of-devops-report https://devops-research.com/ – DORA – DevOps Research & Assessment, authors of the report https://www.wordclouds.com/ https://www.linkedin.com/in/gregwwalters/
Sep 30, 2018
45 min
E031 – Data Collaboration with Cursor
Learn about Cursor, a new platform for collaboration around data, hosted platforms and BI artifacts. I sat down with Adam Weinstein, CEO and Co-Founder of Cursor, to learn about the platform. About Cursor Cursor offers a data search and analytics hub that makes disparate data accessible and actionable, enabling technical and business users alike to effortlessly get answers, collaborate and gain insights. Founded by a trio of data leaders from Salesforce, LinkedIn, and Pandora, Cursor’s easy-to-deploy software has been adopted by teams at Apple, Atlassian, Deloitte, Incedo, LinkedIn, NovumRx, and Slack. Cursor is based in San Francisco, CA. Cursor Press https://techcrunch.com/2018/05/30/cursor-looks-to-build-a-search-tool-for-any-internal-database-with-2m-in-new-funding/ https://www.businessinsider.com/adam-weinstein-linkedin-cursor-2018-7 Topics: What is Adam’s background? How BI has evolved over the past 10-20 years. What some of the most pressing challenges are for organizations today? What should people being doing today, outside of a specific tool, to get better at collaborating? How can Cursor help with those challenges? How is content secured on the platform? (separating data from metadata) Where can people find out more about Cursor? What’s next for Cursor as far as features or a roadmap? What are some tools that Adam can’t live without in your daily work? Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org
Aug 30, 2018
40 min
E030 – July 2018 News Roundup
July 2018 News Roundup This month’s episode is a roundup of news from a variety of sources covering three main topics: BI / Dataviz Tools Databases and Platforms Tools and Frameworks Note: Most of the text extracts below are direct quotations from new sources cited in the source list at the bottom of these show notes. This episode is a compilation from those sources. BI / Dataviz Tools PowerBI enhancements (7/12/18) Microsoft has updated its Power BI analytics service in an effort to expand data prep capabilities and unify data analytics across platforms. “Using the Power Query experience familiar to millions of Power BI Desktop and Excel users, business analysts can ingest, transform, integrate and enrich big data directly in the Power BI web service – including data from a large and growing set of supported on-premises and cloud-based data sources, such as Dynamics 365, Salesforce, Azure SQL Data Warehouse, Excel and SharePoint,” the post reads. Power BI now supports data in Azure Data Lake Storage, and integrates with SQL Server Analysis Services and SQL Server Reporting Services. Microsoft today announced the general availability of Visio Visual for Power BI. Based on the feedback collected from the customers during the preview period, Microsoft has made the following changes to the Visio Visual: Support for Power BI Mobile app The ability to change the diagram link embedded earlier and to copy an embedded link to the clipboard Configurable auto-zoom settings that can be turned on and off Support for complex diagrams using layers Overall performance improvements Tableau acquires Empirical Systems Tableau last month announced the acquisition of Empirical Systems, an artificial intelligence (AI) startup with an automated discovery and analysis engine designed to spot influencers, key drivers, and exceptions in data. Looker Enhances Data Science Capability with Integration for Google Cloud BigQuery ML With Looker and BQML, data teams can now save time and eliminate unnecessary processes by creating machine learning (ML) models directly in Google BigQuery via Looker – without the need to transfer data into additional ML tools. BQML predictive functionality will also be integrated into new or existing Looker Blocks allowing users to surface predictive measures in dashboards and applications. DBs and Platforms MemSQL Unveil Significant Update to Database for Real-time Modern Applications and Analytical Systems (Version 6.5 released) Queries are now up to four times faster than the previous MemSQL version (which was already 10x faster than legacy database providers), enabling insights in milliseconds across billions of rows. New automated workload optimization capabilities provide a consistent database response under ultra-high concurrency without the need for manual tuning or specialized DBA resources. Additions to the MemSQL industry-leading “transform-as-you-ingest” capabilities allow customers to use stored procedures for in-database transformations to easily build real-time data pipelines. Resource optimization improvements for multi-tenant deployments deliver greater control and scalability for varied database sizes whether on-premises or in the cloud. Hortonworks Data Platform 3.0 Even a Hadoop stalwart such as Hortonworks Inc. sees the writing on the wall, which is why, in its recent 3.0 release, it emphasized heterogeneous object storage. The new Hortonworks Data Platform 3.0 supports data storage in all of the major public-cloud object stores, including Amazon S3, Azure Storage Blob, Azure Data Lake, Google Cloud Storage and AWS Elastic MapReduce File System. HDP’s latest storage enhancements include a consistency layer, NameNode enhancements to support scale-out persistence of billions of files with lower storage overhead, and storage-efficiency enhancements such as support for erasure coding across heterogeneous volumes. HDP workloads access non-HDFS cloud storage environments via the Hadoop Compatible File System API. My thoughts: Are Hadoop and HDFS Dying??? As we are heading into the fourth industrial revolution, HDP 3.0 is a giant leap for the Big Data ecosystem, with major changes across the stack and expanded eco-system (Deep Learning and 3rd Party Dockerized Apps). HDP 3.0 can be deployed both on-premise and in the major cloud platforms – AWS, Microsoft Azure, and Google Cloud. Many of the HDP 3.0 new features are based on Apache Hadoop 3.1 and include containerization, GPU support, Erasure Coding and Namenode Federation. In order to provide a Trusted Data Lake, we are installing Apache Ranger and Apache Atlas by default with HDP 3.0. In order to streamline the stack, we have removed components such as Apache Falcon, Apache Mahout, Apache Flume, and Apache Hue, and absorbed Apache Slider functionalities into Apache YARN. Tools and Frameworks Python 3.7.0 is now available Data classes that reduce boilerplate when working with data in classes. A potentially backward-incompatible change involving the handling of exceptions in generators. A “development mode” for the interpreter. Nanosecond-resolution time objects. UTF-8 mode that uses UTF-8 encoding by default in the environment. A new built-in for triggering the debugger. Easier access to debuggers through a new breakpoint() built-in Simple class creation using data classes Customized access to module attributes Improved support for type hinting Higher precision timing functions More importantly, Python 3.7 is fast. Each new release of Python comes with a set of optimizations. In Python 3.7, there are some significant speed-ups, including: There is less overhead in calling many methods in the standard library. Method calls are up to 20% faster in general. The startup time of Python itself is reduced by 10-30%. Importing typing is 7 times faster. You can easily get an idea of how much time the imports in your script takes, using -X importtime: Apache OpenNLP 1.9.0 released The Apache OpenNLP team is pleased to announce the release of Apache OpenNLP 1.9.0. The Apache OpenNLP library is a machine learning based toolkit for the processing of natural language text. It supports the most common NLP tasks, such as tokenization, sentence segmentation, part-of-speech tagging, named entity extraction, chunking, parsing, and coreference resolution. Apache OpenNLP 1.9.0 binary and source distributions are available for download from our download page: download page The OpenNLP library is distributed by Maven Central as well. See the Maven Dependency page for more details: Maven Dependency What’s new in Apache OpenNLP 1.9.0 This release introduces new features, improvements and bug fixes. Java 1.8 and Maven 3.3.9 are required. Additionally the release contains the following changes: Brat Document Parser should support name type filters Brat format support fails on multi fragment annotations Remove MD5 hashes from Release process Use String[] instead of StringList in LanguageModel API BRAT Annotator service Fails to start Token model creation fails without at least one <SPLIT> tag Update Penn Treebank URL Explain the new format of feature generator XML config Unify code to sum up input context features FeatureGeneratorUtil can recognize Japanese Hiragana and Katakana letters TensorFlow 1.9.0 Updated docs for tf.keras: New Keras-based get started and programmers guide page. Update tf.keras to the Keras 2.1.6 API. Added tf.keras.layers.CuDNNGRU and tf.keras.layers.CuDNNLSTM layers. Try it. Adding support of core feature columns and losses to gradient boosted trees estimators. The python interface for the TFLite Optimizing Converter has been expanded, and the command line interface (AKA: toco, tflite_convert) is once again included in the standard pip installation. Improved data-loading and text processing with: tf.decode_compressed tf.string_strip tf.strings.regex_full_match Added experimental support for new pre-made Estimators: tf.contrib.estimator.BaselineEstimator tf.contrib.estimator.RNNClassifier tf.contrib.estimator.RNNEstimator The distributions.Bijector API supports broadcasting for Bijectors with new API changes.   PYPL Language Rankings: Python ranks #1, R at #7 in popularity The new PYPL Popularity of Programming Languages (June 2018) index ranks Python at #1 and R at #7. Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: https://www.techrepublic.com/article/microsoft-power-bi-expansion-aims-to-help-analysts-leverage-business-data-more-easily/ https://mspoweruser.com/visio-visual-for-power-bi-now-available-for-everyone/ https://appsource.microsoft.com/en-us/product/office/WA104381132?src=office&corrid=5ebc6738-bf41-4363-bfc9-d46077451402&omexanonuid=19a90565-f1ed-4800-a00b-8ed084f0fe61 https://www.zdnet.com/article/tableau-takes-next-steps-toward-smart-analytics/ https://insidebigdata.com/2018/07/29/memsql-unveil-significant-update-database-real-time-modern-applications-analytical-systems/ http://blog.revolutionanalytics.com/2018/07/ai-roundup-july-2018.html https://pythoninsider.blogspot.com/2018/06/python-3.html https://www.python.org/downloads/release/python-370/ https://www.infoworld.com/article/3252852/python/whats-new-in-python-37.html https://realpython.com/python37-new-features/ http://blog.revolutionanalytics.com/2018/06/pypl-programming-language-trends.html http://pypl.github.io/PYPL.html https://insidebigdata.com/2018/07/29/looker-enhances-data-science-capability-integration-google-cloud-bigquery-ml/ https://github.com/tensorflow/tensorflow/releases/tag/v1.9.0 https://opennlp.apache.org/news/release-190.html https://siliconangle.com/2018/07/09/hadoops-star-dims-era-cloud-object-data-storage-stream-computing/ https://hortonworks.com/blog/announcing-general-availability-hortonworks-data-platform-3-0-0-ambari-2-7-0-smartsense-1-5-0/
Jul 31, 2018
15 min
E29 – Is Data The New Oil?
Is Data the New Oil? Concept originated by Clive Humby, the British mathematician who established Tesco’s Clubcard loyalty program. Humby highlighted the fact that, although inherently valuable, data needs processing, just as oil needs refining before its true value can be unlocked. Why it is the new oil Valuable commodity Different uses among many applications Currently the big buzz of most large companies (Google, Facebook, Apple, etc.) Quantity is generally better in both AI is the darling of so many industries right now, and it is entirely dependent on data There are ethical concerns with how we source and use this, just like there were and are geopolitical and ethical concerns with how we source and use Oil Certain things cannot function (currently) without oil (passenger airplanes, boats) Same with data: Oil & Gas, Netflix, Agriculture, Manufacturing, Healthcare, a general enabler Why it isn’t the new oil Oil is finite, but data is not Rob’s Counterpoint: There is a shelf life on data that makes it less usable over time Data does not have a standard price benchmark like oil Not a physical asset; can be duplicated or shared relatively easily Oil requires huge amounts of resources to recover and transport Rob’s Counterpoint: building a successful “app” with the scale to generate meaningful data does have some costs, albeit not the scale of oil Data is more useful the more that it is used, whereas oil loses energy the more it is used/processed Rob’s Counterpoint: Oil is not useful by itself to most people; it’s really the product oil becomes or enables that is useful The Data of Oil Difference between operating on surface vs. subsea: small tubing error occurs… Surface: 2-3 hours downtime; a few thousand $$ to fix Subsea: 3 months downtime, $40-50mm to fix, not including lost revenue due to  deferred production (ex. 15,000 bpd well * $67/barrel * 90 days = $90.45mm) A good sized offshore platform generates revenue greater than the entire country of Belize ($2.3bn vs. $1.8bn) Of all the oil we can find, we generally only recover 10-20% in a field with current technology 45-50% of oil generated in the US is used for transportation US consumption per day is about 2 ½ gallons of crude oil / day / person The U.S. has 4% of the world’s population but uses 25% of the world’s oil Total daily oil consumption around the world is 84,249,000 barrels/day Top 3 countries by proven oil reserves are: Venezuela, Saudi Arabia, Canada; US is #10 Gas is 12,200 Wh/kg vs. Li-Ion at 265 Wh/kg (~46x more energy dense) MTTF (Mean time to failure) – 500 years on some parts – needed to operate in subsea environments for 30 years Area of dinner plate = 10.5” Area = Pi * R^2 * 20,000 PSI Area = ¼ * Pi * D^2 * 20,000 PSI 0.25 * pi * 10.5 * 10.5 * 20,000 = 1.73180295029137E6 = 1,731,802 pounds on a single dinner plate (equivalent to ~9 737 Jets) Length Records Analogy: Standing on top of the Empire State Building in NYC and trying to put a straw in a coke can sitting on the sidewalk below Deepest Well (scientific study) = Kola Superdeep Borehole= 40,230 ft. CHAYVO WELL – SAKHALIN-I PROJECT-The current world record holder for longest well; depth of 44,291 feet with a horizontal reach of 39,478 feet DEEPWATER HORIZON – drilled the deepest oil well in history. The well was drilled to 35,050 vertical depth Well depth by year Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: Is Data the New Oil? https://www.economist.com/news/leaders/21721656-data-economy-demands-new-approach-antitrust-rules-worlds-most-valuable-resource https://www.forbes.com/sites/bernardmarr/2018/03/05/heres-why-data-is-not-the-new-oil/#326090363aa9 https://www.forbes.com/sites/howardbaldwin/2014/03/28/big-datas-big-impact-across-industries/#3ebd563856b9 https://hbr.org/2018/01/is-your-companys-data-actually-valuable-in-the-ai-era https://ischool.syr.edu/infospace/2018/02/26/why-data-is-the-new-oil/ The Data of Oil Sources: https://www.bp.com/content/dam/bp/en/corporate/pdf/energy-economics/energy-outlook/bp-energy-outlook-2018.pdf https://oges.info/library/152237/World_s-Deepest-Drilled-Wells http://www.investinganswers.com/investment-ideas/commodities-precious-metals/50-surprising-facts-you-never-knew-about-oil-2692 https://www.dosomething.org/facts/11-facts-about-oil https://oilprice.com/Energy/Crude-Oil/How-Much-Crude-Oil-Do-You-Consume-On-A-Daily-Basis.html https://en.wikipedia.org/wiki/List_of_countries_by_proven_oil_reserves https://xtronics.com/wiki/Energy_density.html https://en.wikipedia.org/wiki/List_of_countries_by_GDP_(nominal) Other Data is the New Oil Sources: https://techcrunch.com/2018/03/27/data-is-not-the-new-oil/amp/ http://ana.blogs.com/maestros/2006/11/data_is_the_new.html https://www.weforum.org/agenda/2018/01/data-is-not-the-new-oil/ https://www.wired.com/insights/2014/07/data-new-oil-digital-economy/ https://www.domo.com/learn/data-never-sleeps-5?aid=DPR072517 https://mobile.cioinsight.com/it-strategy/big-data/data-is-the-new-oil.html https://cloudtweaks.com/2017/01/data-revolution-new-oil/ http://www.oedigital.com/component/k2/item/11383-data-as-the-new-oil http://amp.timeinc.net/fortune/2016/07/11/data-oil-brainstorm-tech http://www.freepressjournal.in/editorspick/ajit-ranade-economic-conclave-data-is-the-new-oil/1108742 Other Oil Sources: https://seekingalpha.com/article/4136900-2018-oil-production-forecast-explained http://www.rrc.state.tx.us/oil-gas/research-and-statistics/ https://www.eia.gov/petroleum/drilling/
Jun 29, 2018
59 min
E028 – Bimodal BI and Data Virtualization
Today we’re back with another guest from the Netherlands. I’m not sure what it is about the Dutch, but they’ve been on a roll with some helpful thought leadership when it comes to data. My guest is Rick van der Lans, a highly-respected analyst, consultant, author, and international lecturer specializing in data warehousing, business intelligence, big data, and database technology. I came across one of Rick’s whitepapers a few months ago on data virtualization. We got in touch and sat down to talk more in depth about the topic. Rick has a lot of data street cred. For many years, he has served as the chairman of the annual European Enterprise Data and Business Intelligence Conference in London and the annual Data Warehousing and Business Intelligence Summit in The Netherlands. He has written tons of articles, blogs, and several books, including the first book on SQL. There will be links to some of the places and things Rick has written and other info in the show notes below. Topics: Rick’s background: author, blogger, consultant – worked on data virtualization (DV) for last 6-7 years How did Rick get interested in DV? Classical data warehouses vs. logical data warehouses What is bi-modal BI? (term introduced by Gartner in 2014) Agile/Self-Service vs. longer, more cautious approach Bi-modal BI vs. the Data Quadrant Comparison of major Data Virtualization Vendors Denodo Tibco DV Manager (bought from Cisco recently) Red Hat Data Virtuality Ultrarep Others (AtScale, Cero, StoneBond, IBM – new entry acquired from Rocket Software) Some are more mature, some are newer (Denodo vs. Tibco = green apples vs. red apples) Companies rolling their own DV (in-memory / views vs. a dedicated tool) DV products are not DB views on steroids Lineage / impact analysis and other features Caching vs. materialization – can store cached data in a virtual table in an intermediary data store. Can be help performance or prevent interference from a transactional source (keeping results consistent for an entire week). How DV can help organizations that are struggling How DV may not be a silver bullet How are different industries embracing these principles? What patterns do you see in companies embracing these principles? What companies should not use this? DV not great at this time on unstructured audio / video, auto-tagging of images Why a classical DWH experienced person may fail at DV What are the warning signs that a DV is going off the rails? Fuzzy logic needed to combine disparate sources Not an integration cureall How you deploy these with projects How to get started? (pick a single, sexy report as a starting point) Where do you go next?  (how to unify other data delivery systems, data marketplaces, API gateways) How to avoid misconceptions about DV (it is slow, only about integration, etc.) How to contact Rick The first book on SQL Places to find Rick’s work: He has published blogs for the following websites: TechTarget.com B-eye-Network Business-analytics.biz He has written the following books: Data Virtualization for Business Intelligence Systems Introduction to SQL (has sold over 100,000 copies and has been translated in many languages) SQL for MySQL Developers The SQL Guide to SQLite The SQL Guide to Ingres The SQL Guide to Pervasive PSQL Music: Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: https://www.denodo.com/en/system/files/document-attachments/rick_van_der_lans_bimodal_ldw.pdf https://solutionsreview.com/data-integration/4-data-virtualization-vendors-to-watch-in-2017/ https://reprints.forrester.com/#/assets/2/545/’RES133042’/reports (Forrester Wave for Data Virtualization) http://www.datavirtualizationblog.com/achieving-lightning-fast-performance-logical-data-warehouse/ http://www.r20.nl/rick_f__van_der_lans1.htm
May 27, 2018
48 min
E027 – The Data Quadrant
My guest in today’s episode is Ronald Damhof (@ronalddamhof), the creator of the Data Quadrant. This quadrant is a sense-making framework in the complex word of data that enables a common frame of reference between managers, domain experts and engineers. This model is used by many organisations to formulate data strategy and justify investments in the data domain. It is used as the strategic underpinning for a data architecture, it guides the ‘rules of the game’ and it separates the fundamental concerns in data. Furthermore, it explains how an organisation can toggle the need to innovate with data and the need to deploy and use data at scale, repeatedly, safe, lawful, with constant quality and robust.     Topics: Ronald background as a “data fundamentalist” His concept of a full scale data architect The push / pull point, from 1950s Toyota, applied to data Development styles from systemic to opportunistic Data Vault’s influence on the quadrant Where data modelling (Q1) and data lakes (Q3) fit into the quadrants Where should you start? Q1/Q2 or Q3/Q4 90% of organizations in the Netherlands are using Data Vault General recommendations on tools by quadrant: Q1 Automation – Wherescape or custom Federalization – mainly still RDBMS Q2 API’ing the data Losing faith in datasets and data marts Q3 Fast infra Doesn’t believe Hadoop is a good fit for most orgs. Likes fast analytical DBs like Vertica or MonetDB? Q4 Open source R, Python, Git, Dataiku Abstraction layer away from code is helpful Azure Platform Ronald Damhof’s background: Primary degree in Economics Certified Data Vault Grand Master Data Architect at the Dutch Central Bank in the Netherlands Music Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: https://twitter.com/ronalddamhof http://prudenza.typepad.com/ (Ronald’s Blog) https://www.linkedin.com/in/ronalddamhof/ (Ronald’s LinkedIn Profile) http://www.tdwi-konferenz.de/tdwi2018/startseite-englisch/program/conference-program/sessiondetails/action/detail/session/di-31-3/title/it-is-all-about-the-data-a-managerial-perspective.html http://www.b-eye-network.com/blogs/damhof/archives/2013/08/4_quadrant_mode.php http://prudenza.typepad.com/dwh/2015/06/make-data-management-a-live-issue-for-discussion-throughout-the-organization.html https://www.linkedin.com/in/thefullscaledataarchitect/
Apr 28, 2018
55 min
E026 – The Four Types of Automation
  Introductory Product Models: Only partner implementations (BluePrism) Limited Features (WorkFusion) Customer Revenue Limited (UI Path) Single License (Softomotive) Music Deep Sky Blue by Graphiqs Groove via FreeMusicArchive.org Sources: https://irpaai.com/definition-and-benefits/ https://www.edgeverve.com/wp-content/uploads/2017/02/forrester-wave-robotic-process-automation.pdf https://www.gartner.com/doc/reprints?id=1-3U26FK2&ct=170222&st=sb http://www.uipath.com/hubfs/News_photos/Forrester_Wave_RPA_Report.png?t=1522186102828 http://images.abbyy.com/India/market_guide_for_robotic_pro_319864%20(002).pdf https://www.uipath.com/community https://www.workfusion.com/rpaexpress https://idm.net.au/article/0011800-which-rpa-software-should-i-use
Mar 30, 2018
27 min
Load more