quinta-feira, 23 de agosto de 2012

Better table search through Machine Learning and Knowledge



The Web offers a trove of structured data in the form of tables. Organizing this collection of information and helping users find the most useful tables is a key mission of Table Search from Google Research. While we are still a long way away from the perfect table search, we made a few steps forward recently by revamping how we determine which tables are "good" (one that contains meaningful structured data) and which ones are "bad" (for example, a table that hold the layout of a Web page). In particular, we switched from a rule-based system to a machine learning classifier that can tease out subtleties from the table features and enables rapid quality improvement iterations. This new classifier is a support vector machine (SVM) that makes use of multiple kernel functions which are automatically combined and optimized using training examples. Several of these kernel combining techniques were in fact studied and developed within Google Research [1,2].

We are also able to achieve a better understanding of the tables by leveraging the Knowledge Graph. In particular, we improved our algorithms for identifying the context and topics of each table, the entities represented in the table and the properties they have. This knowledge not only helps our classifier make a better decision on the quality of the table, but also enables better matching of the table to the user query.

Finally, you will notice that we added an easy way for our users to import Web tables found through Table Search into their Google Drive account as Fusion Tables. Now that we can better identify good tables, the import feature enables our users to further explore the data. Once in Fusion Tables, the data can be visualized, updated, and accessed programmatically using the Fusion Tables API.

These enhancements are just the start. We are continually updating the quality of our Table Search and adding features to it.

Stay tuned for more from Boulos Harb, Afshin Rostamizadeh, Fei Wu, Cong Yu and the rest of the Structured Data Team.


[1] Algorithms for Learning Kernels Based on Centered Alignment
[2] Generalization Bounds for Learning Kernels

quarta-feira, 22 de agosto de 2012

Machine Learning Book for Students and Researchers



Our machine learning book, The Foundations of Machine Learning, is now published! The book, with authors from both Google Research and academia, covers a large variety of fundamental machine learning topics in depth, including the theoretical basis of many learning algorithms and key aspects of their applications. The material presented takes its origin in a machine learning graduate course, "Foundations of Machine Learning", taught by Mehryar Mohri over the past seven years and has considerably benefited from comments and suggestions from students and colleagues at Google.

The book can serve as a textbook for both graduate students and advanced undergraduate students and a reference manual for researchers in machine learning, statistics, and many other related areas. It includes as a supplement introductory material to topics such as linear algebra and optimization and other useful conceptual tools, as well as a large number of exercises at the end of each chapter whose full solutions are provided online.



segunda-feira, 20 de agosto de 2012

Faculty Summit 2012: Online Education Panel



On July 26th, Google's 2012 Faculty Summit hosted computer science professors from around the world for a chance to talk and hear about some of the work done by Google and by our faculty partners. One of the sessions was a panel on Online Education. Daphne Koller's presentation on "Education at Scale" describes how a talk about YouTube at the 2009 Google Faculty Summit was an early inspiration for her, as she was formulating her approach that led to the founding of Coursera. Koller started with the goal of allowing Stanford professors to have more time for meaningful interaction with their students, rather than just lecturing, and ended up with a model based on the flipped classroom, where students watch videos out of class, and then come together to discuss what they have learned. She then refined the flipped classroom to work when there is no classroom, when the interactions occur in online discussion forums rather than in person. She described some fascinating experiments that allow for more flexible types of questions (beyond multiple choice and fill-in-the-blank) by using peer grading of exercises.

In my talk, I describe how I arrived at a similar approach but starting with a different motivation: I wanted a textbook that was more interactive and engaging than a static paper-based book, so I too incorporated short videos and frequent interactions for the Intro to AI class I taught with Sebastian Thrun.

Finally, Bradley Horowitz, Vice President of Product Management for Google+ gave a talk describing the goals of Google+. It is not to build the largest social network; rather it is to understand our users better, so that we can serve them better, while respecting their privacy, and keeping each of their conversations within the appropriate circle of friends. This allows people to have more meaningful conversations, within a limited context, and turns out to be very appropriate to education.

By bringing people together at events like the Faculty Summit, we hope to spark the conversations and ideas that will lead to the next breakthroughs, perhaps in online education, or perhaps in other fields. We'll find out a few years from now what ideas took root at this year's Summit.

sábado, 18 de agosto de 2012

Regular expression matching in C++11


A part of the boost (http://www.boost.org/) functionality regarding regular expressions has been included in the new C++11/C++0x standard.

This code works on GCC and even on Visual C++ 10 (Visual Studio 2010) and above:
    std::regex rgx("(\S+@\S+)");

    std::smatch result;
    std::string str = std::regex_replace(std::string("please send an email to my@mail.com for more information"), rgx, std::string("<$1>"));
    // str contains the same text, but with the e-mail address enclosed between <...>.

More information on regular expressions can be found here: http://www.regular-expressions.info/ .

sexta-feira, 17 de agosto de 2012

Replace all occurrences of a character in a std::string with another character


in one line of C++ code, using C++11/C++0x:

std::string str = "my#string";
std::for_each(str.begin(), str.end(), [] (char &ch) { if (ch == '#') ch = '\\'; }  );
// str now contains my\string

The for_each function code calls the lambda expression (indicated in yellow) for every character, and the lambda expression replaces the '#' with a '\'

terça-feira, 14 de agosto de 2012

The future of technology?

Click the image to make it bigger:


Source: http://envisioningtech.com/

2012 2013 2014 2015 2016 2017 2018 2019 2020 2030 2040 2012 2013 2014 2015 2016 2017 2019 2020 2030 2040 ROBOTICS BIOTECH MATERIALS ENERGY ARTIFICIAL INTELLIGENCE SENSORS GEOENGINEERING QUANTITATIVE FORECASTS INTERNET INTERFACES UBICOMP SPACE BITS ATOMS RELATIVE IMPORTANCE CONSUMER IMPACT CLUSTER OF TECHNOLOGIES The node size indicates the predicted importance of a technology. The outline of a node indicates a consumer impact larger than the technological novelty. A jagged outline indicates a cluster of similar technologies grouped together. World population: 8 billion Source: U.N. – http://bit.ly/7nqQkS World population: 7 billion BRICs GDP overtakes the G7 Source: Goldman Sachs – http://bit.ly/nc9Wqj Petabyte storage standard Source: http://bit.ly/r9BYQc Exabyte storage standard Source: http://bit.ly/kPMKMb Terabit internet speed standard Source: http://bit.ly/kPMKMb World population: 9 billion Source: U.N. – http://bit.ly/7nqQkS Source: http://bit.ly/6MoQJc Sources: Intel – http://intel.ly/pWbH04 Ericsson – http://bit.ly/avvVok Alan Conroy – http://bit.ly/pofHp5 FutureTimeline – http://bit.ly/qz4ben Sources: Intel – http://intel.ly/pWbH04 InternetWorldStats – http://bit.ly/AKbO5 Source: U.N. – http://bit.ly/7nqQkS Global online population: ± 2 billion Connected devices: ±10 billion Global online population: 4-5 billion Connected devices: 30-50 billion $150 Hard disk: ±200 Tb Standard RAM: ±750Gb Global online population: ± 2.5 billion Connected devices: ±15 billion $ 1.000 computer reaches the capacity of the human brain (± 10 15 calculations per second) Vertical farming Weather engineering Seasteading Desalination Carbon sequestration Climate engineering Arcologies Commercial spaceflight Sub-orbital spaceflight Lunar outpost Mars mission Solar sail Space elevator Space tourism Inductive chargers Thorium reactor Traveling wave reactor Fuel cells Multi-segmented smart grids Biomechanical harvesting Bio-enhanced fuels Artificial photosynthesis Space-based solar power Piezoelectricity Photovoltaic glass Nanogenerators Enernet Tidal turbines Programmable matter Personal fabricators Molecular assembler Metamaterials Additive manufacturing Graphene Optical invisibility cloaks Biomaterials Carbon nanotubes Self-healing materials Nanowires Antiaging drugs Stem-cell treatments In-vitro meat Nanomedicine Artificial retinas Rapid personal gene sequencing Synthetic biology Personalized medicine Gene therapy Hybrid assisted limbs Smart drugs Synthetic blood Organ printing Smart toys Robotic surgery Telematics Appliance robots Self-driving vehicles Domestic robots Powered exoskeleton Embodied avatars Swarm robotics Utility fog Commercial UAVs Fabric-embedded screens Reprogrammable chips Picoprojectors Volumetric (3D) screens Flexible screens Skin-embedded screens Modular computers Tablets Boards Retinal screens Eyewear-embedded screens Context-aware computing Smart power meters Biometric sensors Machine vision Optogenetics Depth imaging Biomarkers Neuroinformatics Near-field communication Pervasive video capture Computational photography Speech recognition Haptics 4K Augmented reality Gesture recognition Multi touch Immersive virtual reality Holography Telepresence 4G 5G Cloud computing Interplanetary internet Exocortex Photonics Virtual currencies Cyberwarfare Mesh networking Reputation economy Remote presence VR-only lifeforms Machineaugmented cognition Software agents High-frequency trading Natural language interpretation Procedural storytelling Machine translation Research & visualization by Michell Zappa mz@envisioningtech mz@envisioningtech.com mz@envisioningtech.com Envisioning emerging technology for 2012 and beyond Last updated: 2012-02-10 Understanding where technology is heading is more than guesswork. Looking at emerging trends and research, one can predict and draw conclusions about how the technological sphere is developing, and which technologies should become mainstream in the coming years. Envisioning technology is meant to facilitate these observations by taking a step back and seeing the wider context. By speculating about what lies beyond the horizon we can make better decisions of what to create today. BY SA