terça-feira, 17 de julho de 2012

Reducing PDF file size

Using GhostScript, it is very easy to make PDF files smaller. Run this command:
gs -sDEVICE=pdfwrite -dPDFSETTINGS=/screen -dCompatibilityLevel=1.4 -dNOPAUSE -dQUIET -dBATCH -sOutputFile=NewFile.pdf OriginalFile.pdf

segunda-feira, 16 de julho de 2012

Subversion working copy locked

When you get the SubVersion (svn) error "Working copy <directory> locked" in Tortoise, you can try this to unlock the directory:

  • Open a command prompt, and change the directory to your locked subversion directory.
  • Run svn cleanup

sexta-feira, 13 de julho de 2012

Google at SIGMOD/PODS 2012



Over the years, SIGMOD has expanded beyond a traditional "database" conference to include several areas related to information management. This year’s ACM SIGMOD/PODS conference (on Management of Data, and Principles of Database Systems), held in Scottsdale, Arizona was no different. We were impressed by the wide variety of researchers from industry and academia alike the conference attracted, and enjoyed learning how others are pushing the limits of scalability in data storage and processing. In addition to an excellent set of papers on a large number of topics, we saw a couple of recurring themes:

1) Data Visualization
  • Pat Hanrahan from Stanford gave a keynote on some of the challenges involved in building systems to enable "data enthusiasts" to manage and visualize data. 

2) Big Data


As has been the case for the last couple of years, “Big Data" has been of ever-growing interest to the entire community, particularly from industry. Google presented a talk on F1, a new distributed database system we’ve built to power the AdWords system. A complex business application like AdWords has different requirements than many systems at Google that often use storage systems like Bigtable. We have a single database shared by hundreds of developers and systems, so we need the robustness and ease of use we’re used to from traditional databases. F1 is built to scale like Bigtable, without giving up the database features we also need, like strong consistency, ACID transactions, schema enforcement, and most importantly, SQL query.

There’s been a widespread trend over the last several years away from databases, towards highly scalable “NoSQL” systems. We don’t think that trade-off is necessary, and were happy to see several other speakers advocate a similar theme -- yes, databases are useful, and developers shouldn’t need to give up database features and ease of use in the name of scalability.

This theme was supported by an industry session on Big Data featuring talks from other companies: Facebook (TAO: How Facebook Serves the Social Graph), Twitter (Large-Scale Machine Learning at Twitter), and Microsoft (Recurring Job Optimization in Scope). Googler Kirsten LeFevre was a panelist on the "Perspectives on Big Data" panel organized by Surajit Chaudhuri from Microsoft, and also featuring Donald Kossmann from ETHZ, Sam Madden from MIT, and Anand Rajaraman from Walmart Labs. Last but not the least, Surajit Chaudhuri also gave an excellent keynote outlining some of the research challenges that the new era of "Big Data and Cloud" poses.

As has been the practice for several years now, to continue generating great interest in data management research, SIGMOD has been organizing panels such as this year's "New Research Symposium" (which included Anish Das Sarma from Google as a panelist).

In addition to sponsoring the conference, many Googlers attended contributing to a robust presence and affording us the opportunity to interact with the broader information management community. We've been pushing the frontiers of science with cutting-edge research in many aspects of data management, and we were eager to share our innovations and see what others have been working on. We found Amin Vahdat's keynote on the intersection of Networking and Databases to be a highlight of Google’s participation, which also included presenting papers, participating on panels, and taking part in planning and program committees:

Program Committee Members


Anish Das Sarma, Venkatesh Ganti, Zoltan Gyongyi, Alon Halevy (Tutorials Chair), Kristen LeFevre, Cong Yu

Talks


Symbiosis in Scale Out Networking and Data Management
Amin Vahdat, Google (Keynote)

F1-The Fault-Tolerant Distributed RDBMS Supporting Google's Ad Business
Jeff Shute, Mircea Oancea, Stephan Ellner, Ben Handy, Eric Rollins, Bart Samwel, Radek Vingralek, Chad Whipkey, Xin Chen, Beat Jegerlehner, Kyle Littlefield, Phoenix Tong (Googlers)

Finding Related Tables
Anish Das Sarma, Lujun Fang, Nitin Gupta, Alon Halevy, Hongrae Lee, Fei Wu, Reynold Xin, Cong Yu (Googlers)

Papers


CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster
Changkyu Kim, Jongsoo Park, Nadathur Satish, Hongrae Lee (Google), Pradeep Dubey, Jatin Chhugani

Efficient Spatial Sampling of Large Geographical Tables
Anish Das Sarma, Hongrae Lee, Hector Gonzalez, Jayant Madhavan, Alon Halevy (Googlers)

Panels


Perspectives on Big Data Plenary Session: Privacy and Big Data 
Kristen LeFevre, Google

SIGMOD New Researcher Symposium - How to be a good advisor/advisee? 
Anish Das Sarma, Google

Overall, this year’s SIGMOD was a great conference, widely attended by researchers from industry and academia, and comprised of a very interesting mix of research presentations and discussions. Google had a good showing at the conference, and we look forward to continuing this trend in the coming years.

quinta-feira, 12 de julho de 2012

Reflections on the Google Faculty Institute



Extending the school year one day can result in a year’s worth of learning. This proved true on June 8, when the 2011 Google Faculty Institute (GFI) cohort were welcomed back for a day, to share best practices and perspectives from their funded research over the 2011-12 school year.

For the past year, the GFI Fellows collaborated across 16 California State University campuses, Stanford and UC Berkeley, to execute on ten research initiatives proposed on the final day of the conference. GFI themes of faculty collaboration, project-based learning, universal design and others were implemented in the Fellows’ projects, each of which focused on ways to enhance teaching practices through the use of educational technologies.

At the GFI Redux earlier this month, participants reviewed research initiatives, attended panel discussions, and defined plans for the 2012-13 school year. In a packed day of sessions, the cohort showcased projects ranging from mobile application development to geospatial tool utilization to the success of the flipped classroom. Some highlights of GFI projects:

  • Making Teachers “Appy” presented workshops on UC and CSU campuses on mobile application development using App Inventor. While building confidence with new technologies, participants learned to create their own applications to enhance classroom instruction.
  • Bird’s Eye Detective encouraged CSU pre-service teachers to explore the world from a new perspective utilizing geospatial tools including Google Earth, Google Maps, and Fusion Tables.
  • Transforming STEM Educators included nine hands-on workshops on three CSU campuses, presenting creative ways to engage students in science and engineering courses through the use of technology.
  • CSU Digital Learning Ambassadors are faculty creating collaborative communities and customized initiatives from the inside. Initiatives include tech infusion prizes, Hangouts on Air for academic discussions, and webinars.

The Google Faculty Institute served as a catalyst and incubator for innovative educational technology. Congratulations to the GFI Fellows on a year of excellent research and application.


terça-feira, 3 de julho de 2012

Google Research Awards: Summer, 2012



We’ve just finished the review process for the latest round of the Google Research Awards, which is our bi-annual open call for proposals on research in areas of mutual interest with Google. Our funding provides full-time faculty the opportunity to fund a graduate student and work directly with Google research scientists and engineers.

This round, we are funding 104 awards across 21 different focus areas for a total of nearly $6 million. The subject areas that received the highest level of support this time were systems and infrastructure, human computer interaction, and mobile. In addition, 28% of the funding was awarded to universities outside the U.S.

Given that our program is merit-based, we make funding decisions via committees of experts, who assess each proposal by its impact, innovation, relevance to Google, and other factors. Over the past two years, we have seen significant growth in the Research Award program. This round, we had 815 proposals—up 11% from last round, which required 1,946 reviews by 654 reviewers.

Our award committees represent a microcosm of Research @ Google. Not only do we work with research scientists in making funding decisions, but also engineers—many of whom have advanced degrees in Computer Science. Our research organization has a similar make-up: both research scientists and engineers working together on innovative projects that are product-focused and relevant to our customers.

Congratulations to the well-deserving recipients of this round’s awards. If you are interested in applying for the next round (deadline is October 15), please visit our website for more information.

The following packages have been kept back

When executing "sudo apt-get upgrade" on a Ubuntu server, you can get the message:
The following packages have been kept back

This means that certain updates require system changes.  You can install the packages by executing this command:
sudo apt-get dist-upgrade

segunda-feira, 2 de julho de 2012

Our Unique Approach to Research



Google started as a research project—and research has remained a core part of our culture. But we also do research differently than many other places. To shed more light on Google’s unique approach to research, Peter Norvig (Director of Research), Slav Petrov (Senior Research Scientist) and I recently published a paper, “Google’s Hybrid Approach to Research,” in the July issue of Communications of the ACM.
   
In the paper, we describe our hybrid approach to research, which integrates research and development to maximize our impact on users and the speed at which we make progress. Our model allows us to work at unparalleled scale and conduct research in vivo on real systems with millions of users, rather than on artificial prototypes. This yields not only innovative research results and new technologies, but valuable new capabilities for the company—think of MapReduce, Voice Search or open source projects such as Android and Chrome. 

Breaking up long-term research projects into shorter-term, measurable components is another aspect of our integrated model. This is not to say our model precludes longer-term objectives, but we try to achieve these in stages. For example, Google Translate is a multi-year project characterized by the need for both research and complex systems, but we’ve achieved many small objectives along the way—such as adding languages over time for a current total of 64, developing features like two-step translation functionality, enabling users to make corrections, and consideration of syntactic structure.

Overall, our success in the areas of systems, speech recognition, language translation, machine learning, market algorithms, computer vision and many other areas has stemmed from our hybrid research approach. While there are risks associated with the close integration of research and development activities—namely the concern that research will take a back seat in favor of shorter-term projects—we mitigate those by focusing on the user and empirical data, maintaining a flexible organizational structure, and engaging with the academic community. We have a portfolio of timescales, with some researchers working with engineers to rapidly iterate on existing products, and others working on forward-looking projects that will benefit people in the future.

We hope “Google’s Hybrid Approach to Research” helps explain our method. We feel it will bring some clarification and transparency to our approach, and perhaps merit consideration by other technology companies and academic labs that organize research differently.

To learn more about what we do and see see real-time applications of our hybrid research model, add Research at Google to your circles on Google+.

(Cross-posted on the Official Google Blog