How do you like this for the future of Journalism??:Digital Domain
When the Software Is the Sportswriter
By RANDALL STROSS
Published: November 27, 2010
ONLY human writers can distill a heap of sports statistics into a compelling story. Or so we human writers like to think.
The StatSheet site for Ohio State men's basketball. One recent story on the site had 10 sentences and 156 words — written entirely by software.
StatSheet, a Durham, N.C., company that serves up sports statistics in monster-size portions, thinks otherwise. The company, with nine employees, is working to endow software with the ability to turn game statistics into articles about college basketball games.
Established in 2007, StatSheet.com provides statistical analysis of college football and basketball, Nascar and other sports. It dices data in more ways than any fan could possibly absorb. But charts, graphs and rankings alone cannot replace words that tell a story. We humans love stories; a craving for narrative seems part of our nature.
This month, StatSheet unveiled StatSheet Network, made up of separate Web sites for each of the 345 N.C.A.A. Division I men’s basketball teams. Beyond statistics galore, each site has what the company calls “automated content,” stories written entirely by software, including write-ups of the team’s games, past and future. With a joking wink, StatSheet’s founder, Robbie Allen, refers to these sites as the “Robot Army.”
Each team’s StatSheet Web site is located at a freestanding Web address, conveying the sense that it is wholly invested in the interests of that school’s fans. (To find a domain name, a fan first visits http://statsheet.com/#websites.)
The software is imbued with the smarts to flatter each particular team. The same statistics, documenting the same game, produce an entirely different write-up and headline at the opposing team’s page.
A team like No. 1-ranked Duke — whose StatSheet Network Web site is at BlueDevilDaily.com — does not lack for attention from human sports writers. But StatSheet expects that the sports programs of smaller schools will appreciate the advent of robot journalism.
“There are at least 200 Division I schools that the large sports media companies give no attention to,” says Mr. Allen at StatSheet. “Once we have the algorithm in place, there’s no cost to adding the Lamars and Elons to the Dukes and U.N.C.’s.”
Small schools are less likely to have large alumni bases and to draw significant traffic, Mr. Allen said, so he is knocking on their doors to explore licensing partnerships.
Mr. Allen explains that his story-writing software does not perform linguistic analysis; it just uses template sentences and a database of phrases that numbers about 5,000 for now.
“My goal was that 80 percent of readers wouldn’t question that the content was written by a human,” he says, “and now that we’ve launched, I think the percentage is higher.”
In the battle between human writers and robots, I’m not a dispassionate observer: I own up to rooting for the home team. So I asked Michael W. White, an assistant professor of linguistics at Ohio State and a specialist in the field of natural language generation, for his professional opinion of the quality of the robot army’s writing.
We visited the StatSheet site for Ohio State’s team —BuckeyesBeat.com — and looked at the write-up of its Nov. 12 season opener, written expressly for Buckeyes fans: ”Ohio State Gets 102-61 Monster Win Over North Carolina A&T.” “Ohio State has already started living up to monumental expectations with a good first game,” it began. “On November 12th on their home court, the Buckeyes waxed the Aggies, 102-61.”
The story had 10 sentences and 156 words. Over all, Professor White said, it read “pretty well.” He praised the first sentence as very good and said the use of “waxed” in the second was a nice touch.
Then he pointed out some glitches. In one passage, the software lost track of who had beaten whom and forgot to use “the” when referring to Ohio State’s supposed victory “over Buckeyes.” A note that followed the write-up was a bit overeager to show that Ohio State was undefeated so far this season, when its record was still just 1-0.
The Statsheet Network’s Web sites are plainly labeled “beta,” and these minor bugs can easily be eliminated with tweaks. The bigger problem, the professor said, is that StatSheet’s software lacks the awareness of linguistic structures that would let it generate more complex sentences: “Getting from remarkably good to human level is very hard. The more you cram into a sentence, the more difficult it is to do it well.”
The StatSheet software avoids those difficulties by using simple sentences and swapping in particular details. “This makes it perfectly readable, but it also can make it seem stilted,” he said.
Mr. Allen of StatSheet says he believes that what some readers regard as “stilted” will be appreciated by others who say “‘ I don’t like personality — I just want the straight facts.’”
He sees opportunities to extend robo-journalism not only to other sports but also to other subject areas — like financial news — in which there is an abundance of readily available data.
Will StatSheet’s robo journalists replace business columnists? Not a chance: we humans are undefeated.
Only problem is, just like the Buckeyes two weeks ago, we’re just 1-0.
Randall Stross is an author based in Silicon Valley and a professor of business at San Jose State University. E-mail: stross@nytimes.com.
By JEREMY W. PETERS
Published: September 5, 2010
Daniel Rosenbaum for The New York Times
The Washington Post newsroom displays traffic data. Raju Narisetti, a managing editor, said it helped to decide where to cut staff.
Now, because of technology that can pinpoint what people online are viewing and commenting on, how much time they spend with an article and even how much money an article makes in advertising revenue, newspapers can make more scientific decisions about allocating their ever scarcer resources.
Such data has never been available with such specificity and timeliness. The reader surveys that newspapers relied on for decades took months to produce, often leaving editors with stale data.
Looking to the public for insight on how to cover a topic is never comfortable for newsrooms, which have the deeply held belief that readers come to a newspaper not only for its information but also for its editorial judgment. But many newsrooms now seem to be re-examining that idea and embracing, albeit cautiously, a more democratic approach to serving up the news, particularly online.
“How can you say you don’t care what your customers think?” asked Alan Murray, who oversees online news at The Wall Street Journal. “We care a lot about what our readers think. But our readers also care a lot about our editorial judgment. So we’re always trying to balance the two.”
Editors at The Journal, like those at other large newspapers, follow the Web traffic metrics closely. The paper’s top editors begin their morning news meetings with a rundown of data points, including the most popular search terms on WSJ.com, which articles are generating the most traffic and what posts are generating buzz on Twitter.
At The Washington Post, a television screen with an array of data — the number of unique visitors to washingtonpost.com, how many articles those visitors view and where on the Web those visitors came from — is on display for the entire newsroom. A red or green marker designates each data point, indicating whether the Web site’s goal for the month on that particular metric has been met. About 120 people in The Post’s newsroom get an e-mail each day laying out how the Web site performed in the closely watched metrics — 46 in all.
Rather than corrupt news judgment by causing editors to pander to the most base reader interests, the availability of this technology so far seems to be leading to more surgical decisions about how to cover a topic so it becomes more appealing to an online audience.
The Post, which provided extensive coverage of the recent elections in Britain online and in its print editions, found that online readers were not particularly interested in the topic. One of the five most viewed items on The Post’s Web site in the last year, in fact, was not a political project at all but a piece on Crocs, the popular foam footwear. Editors attributed that to Yahoo, which linked to the article.
But that did not translate into more Croc coverage. And coverage of the British elections was not scaled back.
Raju Narisetti, The Post’s managing editor overseeing online operations, said he saw reader metrics as a tool to help him better determine how to use online resources.
“We ask, ‘What can we do online to make it more attractive?” ’ Mr. Narisetti said. “Can we do podcasts? Can we do a photo gallery? Can we do any kind of user-generated content?”
He said the data has proved highly useful in today’s world of shrinking newsroom budgets. Mr. Narisetti said that when he had to reduce his staff last year, he looked at what kind of content was not performing well with readers. He discovered that long-form video had a low audience, so he reduced that department by a couple of people.
At The Journal, editors use traffic data to inform decisions on how articles should be presented on WSJ.com. “We look at the data, and if things are getting a lot of hits, they’ll get better play and longer play on the home page,” said Mr. Murray. Conversely, articles getting low audiences will be moved down more quickly if there is no compelling news reason to keep them prominent.
But Mr. Murray explained that the data was not always used as a blunt tool. In the case of a rather dry business development last month involving the Potash Corporation, the Canadian fertilizer maker, Journal editors decided to prominently display articles on the subject despite very low traffic numbers.
“We didn’t put it there because it was going to be a big traffic getter. We put it there because it’s big important news in the business world,” Mr. Murray said.
The New York Times does not use Web metrics to determine how articles are presented, but it does use them to make strategic decisions about its online report, said Bill Keller, the executive editor. “We don’t let metrics dictate our assignments and play,” he said, “because we believe readers come to us for our judgment, not the judgment of the crowd. We’re not ‘American Idol.’ ”
Mr. Keller added that the paper would, for example, use the data to determine which blogs to expand, eliminate or tweak.
As newspaper Web sites use technology to learn more about readers’ habits, they are also developing new ways to persuade readers to tell them more about what they want. The Los Angeles Times features what it calls a “personality quiz” for readers on its Web site. The feature adds a spin to the personalization options that Web sites have offered for the last few years with a 17-question test that asks readers things like “What does success mean to you?” and has them pick from 12 photos. A few options include images of a wedding, a gleaming sports car and a man embracing a peasant child.
At the end of the quiz, readers are assigned a personality type like “dynamo,” who, as the quiz explains, is someone “always seeking new adventures that broaden your horizons and take you out of your comfort zone.” A customized news feed then appears each time a reader visits the Web site from the same computer.
“It helps me understand the readers in a way that I can’t with just the metrics,” said Sean Gallagher, managing editor for online operations at The Los Angeles Times, explaining that he now pairs sports articles with food articles because surveys have shown a correlation.
As the technology advances and allows papers to look more deeply at performance metrics, newsrooms may find that there is just some data they would rather not know.
At a recent meeting with the top online editors of The Los Angeles Times, a consulting group that helps media companies enhance profits from their Web sites pitched new software that it said could change the industry. The newsroom would be able to know how much money — down to the penny — each of its articles online was making when readers clicked on ads.
“I could see a business case for it,” said Mr. Gallagher, who hastened to add, “I don’t agree with that business case.”
Software developers acknowledge that the questions can be difficult as newspapers try to reinvent their business models. But they say the dialogue is ultimately constructive.
“By having this data and making it available, we’re spurring the conversations to take place,” said Tim Ruder, chief revenue officer for Perfect Market, the company that developed the tracking software for ad clicks. “And it’s especially healthy to have those conversations in the context of experience and not in an abstract way.”
Such data has never been available with such specificity and timeliness. The reader surveys that newspapers relied on for decades took months to produce, often leaving editors with stale data.
Looking to the public for insight on how to cover a topic is never comfortable for newsrooms, which have the deeply held belief that readers come to a newspaper not only for its information but also for its editorial judgment. But many newsrooms now seem to be re-examining that idea and embracing, albeit cautiously, a more democratic approach to serving up the news, particularly online.
“How can you say you don’t care what your customers think?” asked Alan Murray, who oversees online news at The Wall Street Journal. “We care a lot about what our readers think. But our readers also care a lot about our editorial judgment. So we’re always trying to balance the two.”
Editors at The Journal, like those at other large newspapers, follow the Web traffic metrics closely. The paper’s top editors begin their morning news meetings with a rundown of data points, including the most popular search terms on WSJ.com, which articles are generating the most traffic and what posts are generating buzz on Twitter.
At The Washington Post, a television screen with an array of data — the number of unique visitors to washingtonpost.com, how many articles those visitors view and where on the Web those visitors came from — is on display for the entire newsroom. A red or green marker designates each data point, indicating whether the Web site’s goal for the month on that particular metric has been met. About 120 people in The Post’s newsroom get an e-mail each day laying out how the Web site performed in the closely watched metrics — 46 in all.
Rather than corrupt news judgment by causing editors to pander to the most base reader interests, the availability of this technology so far seems to be leading to more surgical decisions about how to cover a topic so it becomes more appealing to an online audience.
The Post, which provided extensive coverage of the recent elections in Britain online and in its print editions, found that online readers were not particularly interested in the topic. One of the five most viewed items on The Post’s Web site in the last year, in fact, was not a political project at all but a piece on Crocs, the popular foam footwear. Editors attributed that to Yahoo, which linked to the article.
But that did not translate into more Croc coverage. And coverage of the British elections was not scaled back.
Raju Narisetti, The Post’s managing editor overseeing online operations, said he saw reader metrics as a tool to help him better determine how to use online resources.
“We ask, ‘What can we do online to make it more attractive?” ’ Mr. Narisetti said. “Can we do podcasts? Can we do a photo gallery? Can we do any kind of user-generated content?”
He said the data has proved highly useful in today’s world of shrinking newsroom budgets. Mr. Narisetti said that when he had to reduce his staff last year, he looked at what kind of content was not performing well with readers. He discovered that long-form video had a low audience, so he reduced that department by a couple of people.
At The Journal, editors use traffic data to inform decisions on how articles should be presented on WSJ.com. “We look at the data, and if things are getting a lot of hits, they’ll get better play and longer play on the home page,” said Mr. Murray. Conversely, articles getting low audiences will be moved down more quickly if there is no compelling news reason to keep them prominent.
But Mr. Murray explained that the data was not always used as a blunt tool. In the case of a rather dry business development last month involving the Potash Corporation, the Canadian fertilizer maker, Journal editors decided to prominently display articles on the subject despite very low traffic numbers.
“We didn’t put it there because it was going to be a big traffic getter. We put it there because it’s big important news in the business world,” Mr. Murray said.
The New York Times does not use Web metrics to determine how articles are presented, but it does use them to make strategic decisions about its online report, said Bill Keller, the executive editor. “We don’t let metrics dictate our assignments and play,” he said, “because we believe readers come to us for our judgment, not the judgment of the crowd. We’re not ‘American Idol.’ ”
Mr. Keller added that the paper would, for example, use the data to determine which blogs to expand, eliminate or tweak.
As newspaper Web sites use technology to learn more about readers’ habits, they are also developing new ways to persuade readers to tell them more about what they want. The Los Angeles Times features what it calls a “personality quiz” for readers on its Web site. The feature adds a spin to the personalization options that Web sites have offered for the last few years with a 17-question test that asks readers things like “What does success mean to you?” and has them pick from 12 photos. A few options include images of a wedding, a gleaming sports car and a man embracing a peasant child.
At the end of the quiz, readers are assigned a personality type like “dynamo,” who, as the quiz explains, is someone “always seeking new adventures that broaden your horizons and take you out of your comfort zone.” A customized news feed then appears each time a reader visits the Web site from the same computer.
“It helps me understand the readers in a way that I can’t with just the metrics,” said Sean Gallagher, managing editor for online operations at The Los Angeles Times, explaining that he now pairs sports articles with food articles because surveys have shown a correlation.
As the technology advances and allows papers to look more deeply at performance metrics, newsrooms may find that there is just some data they would rather not know.
At a recent meeting with the top online editors of The Los Angeles Times, a consulting group that helps media companies enhance profits from their Web sites pitched new software that it said could change the industry. The newsroom would be able to know how much money — down to the penny — each of its articles online was making when readers clicked on ads.
“I could see a business case for it,” said Mr. Gallagher, who hastened to add, “I don’t agree with that business case.”
Software developers acknowledge that the questions can be difficult as newspapers try to reinvent their business models. But they say the dialogue is ultimately constructive.
“By having this data and making it available, we’re spurring the conversations to take place,” said Tim Ruder, chief revenue officer for Perfect Market, the company that developed the tracking software for ad clicks. “And it’s especially healthy to have those conversations in the context of experience and not in an abstract way.”