segunda-feira, 17 de dezembro de 2012

Millions of Core-Hours Awarded to Science



In 2011 Google University Relations launched a new academic research awards program, Google Exacycle for Visiting Faculty, offering up to one billion core-hours to qualifying proposals. We were looking for projects that would consume 100M+ core-hours each and be of critical benefit to society. Not surprisingly, there was no shortage of applications.

Since then, the following seven scientists have been working on-site at Google offices in Mountain View and Seattle. They are here to run large computing experiments on Google’s infrastructure to change the future. Their projects include exploring antibiotic drug resistance, protein folding and structural modelling, drug discovery, and last but not least, the dynamic universe.

Today, we would like to introduce the Exacycle award recipients and their work. Please stay tuned for updates next year.

Simulating a Dynamic Universe with the Large Synoptic Sky Survey
Jeff Gardner, University of Washington, Seattle, WA
Collaborators: Andrew Connolly, University of Washington, Seattle, WA, and John Peterson, Purdue University, West Lafayette, IN

Research subject: The Large Synoptic Survey Telescope (LSST) is one of the most ambitious astrophysical research programs ever undertaken. Starting in 2019, the LSST’s 3.2 Gigapixel camera will repeatedly survey the southern sky, generating tens of petabytes of data every year. The images and catalogs from the LSST have the potential to transform both our understanding of the universe and the way that we engage in science in general.
Exacycle impact: In order to design the telescope to yield the best possible science, the LSST collaboration has undertaken a formidable computational campaign to simulate the telescope itself. This will optimize how the LSST surveys the sky and provide realistic datasets for the development of analysis pipelines that can operate on hundreds of petabytes. Using Exacycle, we are reducing the time required to simulate one night of LSST observing, roughly 5 million images, from 3 months down to a few days. This rapid turnaround will enable the LSST engineering teams to test new designs and new algorithms with unprecedented precision, which will ultimately lead to bigger and better science from the LSST.

Designing and Defeating Antibiotic Drug Resistance
Peter Kasson, Assistant Professor, Departments of Molecular Physiology and Biological Physics and of Biomedical Engineering, University of Virginia

Research subject: Antibiotics have made most bacterial infections routinely treatable. As antibiotic use has become common, bacterial resistance to these drugs has also increased. Recently, some bacteria have arisen that are resistant to almost all antibiotics. We are studying the basis for this resistance, in particular the enzyme that acts to break down many antibiotics. Identifying the critical changes required for pan-resistance will aid surveillance and prevention; it will also help elucidate targets for the development of new therapeutic agents.
Exacycle impact: Exacycle allows us to simulate the structure and dynamics of several thousand enzyme variants in great detail. The structural differences between enzymes from resistant and non-resistant bacteria are subtle, so we have developed methods to compare structural "fingerprints" of the enzymes and identify distinguishing characteristics. The complexity of this calculation and large number of potential bacterial sequences mean that this is a computationally intensive task; the massive computing power offered by Exacycle in combination with some novel sampling strategies make this calculation tractable.


Sampling the conformational space of G protein-coupled receptors
Kai Kohlhoff, Research Scientist at Google
Collaborators: Research labs of Vijay Pande and Russ Altman at Stanford University

Research subject: G protein-coupled receptors (GPCRs) are proteins that act as signal transducers in the cell membrane and influence the response of a cell to a variety of external stimuli. GPCRs play a role in many human diseases, such as asthma and hypertension, and are well established as a primary drug target.
Exacycle impact: Exacycle let us perform many tens of thousands of molecular simulations of membrane-bound GPCRs in parallel using the Gromacs software. With MapReduce, Dremel, and other technologies, we analyzed the 100s of Terabytes of generated data and built Markov State Models. The information contained in these models can help scientists design drugs that have higher potency and specificity than those presently available.
Results: Our models let us explore kinetically meaningful receptor states and transition rates, which improved our understanding of the structural changes that take place during activation of a signaling receptor. In addition, we used Exacycle to study the affinity of drug molecules when binding to different receptor states.


Modeling transport through the nuclear pore complex
Daniel Russel, post doc in structural biology, University of California, San Francisco

Research subject: Our goal is to develop a predictive model of transport through the nuclear pore complex (NPC). Developing the model requires understanding how the behavior of the NPC varies as we change the parameters governing the components of the system. Such a model will allow us to understand how transportins, the unstructured domains and the rest of the cellular milieu, interact to determine efficiency and specificity of macromolecular transport into and out of the nucleus.
Exacycle impact: Since data describing the microscopic behavior of most parts of the nuclear transport process is incomplete and contradictory, we have to explore a larger parameter space than would be feasible with traditional computational resources.
Status: We are currently modeling various experimental measurements of aspects of the nuclear transport process. These experiments range from simple ones containing only a few components of the transport process to measurements on the whole nuclear pore with transportins and cellular milieu.


Large scale screening for new drug leads that modulate the activity of disease-relevant proteins
James Swetnam, Scientific Software Engineer, drugable.org, NYU School of Medicine
Collaborators: Tim Cardozo, MD, PhD - NYU School of Medicine.

Research subject: We are using a high throughput, CPU-bound procedure known as virtual ligand screening to ‘dock’, or produce rough estimates of binding energy, for a large sample of bioactive chemical space to the entirety of known protein structures. Our goal is the first computational picture of how bioactive chemistry with therapeutic potential can affect human and pathogen biology.
Exacycle Impact: Typically, using our academic lab’s resources, we could screen a few tens of thousands of compounds against a single protein to try to find modulators of its function. To date, Exacycle has enabled us to screen 545,130 compounds against 8,535 protein structures that are involved in important and underserved diseases as cancer, diabetes, malaria, and HIV to look for new leads towards future drugs.
Status: We are currently expanding our screens to an additional 206,190 models from
ModBase. We aim to have a public dataset for the research community in the first half of 2013.

Protein Structure Prediction and Design
Michael Tyka, Research Fellow, University of Washington, Seattle, WA

Research subject: The precise relationship between the primary sequence and the three dimensional structure of proteins is one of the unsolved grand challenges of computational biochemistry. The Baker Lab has made significant progress in recent years by developing more powerful protein prediction and design algorithms using the Rosetta Protein Modelling suite.
Exacycle impact: Limitations in the accuracy of the physical model and lack of sufficient computational power have prevented solutions to broader classes of medically relevant problems. Exacycle allows us to improve model quality by conducting large parameter optimization sweeps with a very large dataset of experimental protein structural data. The improved energy functions will benefit the entire theoretical protein research community.

We are also using Exacycle to conduct simultaneous docking and one-sided protein design to develop novel protein binders for a number of medically relevant targets. For the first time, we are able to aggressively redesign backbone conformations at the binding site. This allows for a much greater flexibility in possible binding shapes but also hugely increases the space of possibilities that have to be sampled. Very promising designs have already been found using this method.

quinta-feira, 13 de dezembro de 2012

Continuing the quest for future computer scientists with CS4HS



Computer Science for High School (CS4HS) began five years ago with a simple question: How can we help create a much needed influx of CS majors into universities and the workforce? We took our questions to three of our university partners--University of Washington, Carnegie Mellon, and UCLA--and together we came up with CS4HS. The model was based on a “train the trainer” technique. By focusing our efforts on teachers and bringing them the skills they need to implement CS into their classrooms, we would be able to reach even more students. With grants from Google, our partner universities created curriculum and put together hands-on, community-based workshops for their local area teachers.

Since the initial experiment, CS4HS has exploded into a worldwide program, reaching more than 4,000 teachers and 200,000 students either directly or indirectly in more than 34 countries. These hands-on, in-person workshops are a hallmark of our program, and we will continue to fund these projects going forward. (For information on how to apply, please see our website.) The success of this popular program speaks for itself, as we receive more quality proposals each year. But success comes at a price, and we have found that the current format of the workshops is not infinitely scalable.

This is where Research at Google comes in. This year, we are experimenting with a new model for CS4HS workshops. By harnessing the success of online courses such as Power Searching with Google, and utilizing open-source platforms like the one found in Course Builder, we are hoping to put the “M” in “MOOC” and reach a broader audience of educators, eager to learn how to teach CS in their classrooms.

For this pilot, we are looking to sponsor two online workshops, one that is geared toward CS teachers, and one that is geared toward CS for non-CS teachers to go live in 2013. This is a way for a university (or several colleges working together) to create one incredible workshop that has the potential to reach thousands of enthusiastic teachers. Just as with our in-person workshops, applications will be open to college, university, and technical schools of higher learning only, as we depend on their curriculum expertise to put together the most engaging programs. For this pilot, we will be looking for MOOC proposals in the US and Canada only.

We are really excited about the possibilities of this new format, and we are looking for quality applications to fund. While applications don’t have to run on our Course Builder platform, we may be able to offer some additional support to funded projects that do. If you are interested in joining our experiment or just learning more, you can find information on how to apply on our CS4HS website (or click here).

Applications are open until February 16, 2013; we can’t wait to see what you come up with. If you have questions, please email us at cs4hs@google.com.

quinta-feira, 8 de novembro de 2012

Desenvolvimento de Sites

Procura um serviço de qualidade, para desenvolver seu site ou o site da sua empresa? 

Conheça nosso portfólio e faça seu orçamento agora mesmo!

Desenvolvimento de sites de alto padrão, tudo o que você ou sua empresa precisam! Entre em contato agora mesmo e faça seu orçamento!

Exemplo de sites criados:


quarta-feira, 31 de outubro de 2012

Large Scale Language Modeling in Automatic Speech Recognition



At Google, we’re able to use the large amounts of data made available by the Web’s fast growth. Two such data sources are the anonymized queries on google.com and the web itself. They help improve automatic speech recognition through large language models: Voice Search makes use of the former, whereas YouTube speech transcription benefits significantly from the latter.

The language model is the component of a speech recognizer that assigns a probability to the next word in a sentence given the previous ones. As an example, if the previous words are “new york”, the model would assign a higher probability to “pizza” than say “granola”. The n-gram approach to language modeling (predicting the next word based on the previous n-1 words) is particularly well-suited to such large amounts of data: it scales gracefully, and the non-parametric nature of the model allows it to grow with more data. For example, on Voice Search we were able to train and evaluate 5-gram language models consisting of 12 billion n-grams, built using large vocabularies (1 million words), and trained on as many as 230 billion words.


The computational effort pays off, as highlighted by the plot above: both word error rate (a measure of speech recognition accuracy) and search error rate (a metric we use to evaluate the output of the speech recognition system when used in a search engine) decrease significantly with larger language models.

A more detailed summary of results on Voice Search and a few YouTube speech transcription tasks (authors: Ciprian Chelba, Dan Bikel, Maria Shugrina, Patrick Nguyen, Shankar Kumar) presents our results when increasing both the amount of training data, and the size of the language model estimated from such data. Depending on the task, availability and amount of training data used, as well as language model size and the performance of the underlying speech recognizer, we observe reductions in word error rate between 6% and 10% relative, for systems on a wide range of operating points.

domingo, 28 de outubro de 2012

Ubuntu Server distribution upgrade

Doing a distribution upgrade is easy in the graphical user interface.  On a "headless" Ubuntu Server, you probably don't have a graphical user interface available, and you need to perform a few easy steps.

First you need to check 2 things:
  • The package "update-manager-core" needs to be installed.  Just run the following command: (it will install it, if it is not already)
    sudo apt-get install update-manager-core
  • On Ubuntu Server, the update channel is usually "LTS" (Long Time Support) which will upgrade only to the major release of Ubuntu Server.  Probably, you want to set this to "normal", to allow an upgrade to every distribution upgrade.  To do this open the file "/etc/update-manager/release-upgrades", and change the line "Prompt=lts" to "Prompt=normal".
    For example, with pico:
    sudo pico /etc/update-manager/release-upgrades

    Change the Prompt setting to "normal".

    Press Ctrl-X, press "Y" to confirm and press enter to confirm the filename.

Then the distribution upgrade can be started by running this command:
sudo do-release-upgrade -d 

It is not "recommended" to do a distribution upgrade remotely over ssh.  The "do-release-upgrade" command above checks on this, and gives a warning.  However, I didn't have any problem doing the upgrade remotely over ssh.

quinta-feira, 18 de outubro de 2012

Converting a date/time to time_t using FxGqlC (or SQL)

The time_t is used in C++ to represent a date/time.  It is expressed in seconds since Januari 1st, 1970.

To get the current date/time as a time_t value, you can run this query in FxGqlC (or SQL):
select datediff(second, '1970-01-01', getutcdate())

You need to use getutcdate() because time_t defines the UTC time.

Or for any arbitrary date/time (in UTC):
select datediff(second, '1970-01-01', '2012-10-18 22:33') 
-- Returns 1350599580

The other way around is also easy: run this query to convert a time_t to a date/time
select dateadd(second, 1234567890, '1970-01-01')
-- Returns 13/02/2009 23:31:30

Ngram Viewer 2.0



Since launching the Google Books Ngram Viewer, we’ve been overjoyed by the public reception. Co-creator Will Brockman and I hoped that the ability to track the usage of phrases across time would be of interest to professional linguists, historians, and bibliophiles. What we didn’t expect was its popularity among casual users. Since the launch in 2010, the Ngram Viewer has been used about 50 times every minute to explore how phrases have been used in books spanning the centuries. That’s over 45 million graphs created, each one a glimpse into the history of the written word. For instance, comparing flapper, hippie, and yuppie, you can see when each word peaked:

Meanwhile, Google Books reached a milestone, having scanned 20 million books. That’s approximately one-seventh of all the books published since Gutenberg invented the printing press. We’ve updated the Ngram Viewer datasets to include a lot of those new books we’ve scanned, as well as improvements our engineers made in OCR and in hammering out inconsistencies between library and publisher metadata. (We’ve kept the old dataset around for scientists pursuing empirical, replicable language experiments such as the ones Jean-Baptiste Michel and Erez Lieberman Aiden conducted for our Science paper.)

At Google, we’re also trying to understand the meaning behind what people write, and to do that it helps to understand grammar. Last summer Slav Petrov of Google’s Natural Language Processing group and his intern Yuri Lin (who’s since joined Google full-time) built a system that identified parts of speech—nouns, adverbs, conjunctions and so forth—for all of the words in the millions of Ngram Viewer books. Now, for instance, you can compare the verb and noun forms of “cheer” to see how the frequencies have converged over time:
Some users requested the ability to combine Ngrams, and Googler Matthew Gray generalized that notion into what we’re calling Ngram compositions: the ability to add, subtract, multiply, and divide Ngram counts. For instance, you can see how “record player” rose at the expense of “Victrola”:
Our info page explains all the details about this curious notion of treating phrases like components of a mathematical expression. We’re guessing they’ll only be of interest to lexicographers, but then again that’s what we thought about Ngram Viewer 1.0.

Oh, and we added Italian too, supplementing our current languages: English, Chinese, Spanish, French, German, Hebrew, and Russian. Buon divertimento!