Showing posts with label text analysis. Show all posts
Showing posts with label text analysis. Show all posts

Wednesday, October 4, 2017

An Adventure in Google Ngram

In these next few weeks, my plan is to use my blog posts to explore programs that I am considering using in my thesis work and, since I'm already familiar with Voyant and Hypothes.is, I'm going to start this week by researching Google Ngram.

In one conversation between Alan and myself, the topic of accessibility came up, and we tossed around the issue of how to determine if a tool is reasonably easy to learn. Between myself and my target group, few DH newcomers are going to have the time or expertise to learn the more complicated programs. As I discover tools that I'm interested in using, I'm going to have to come with with an accessibility scale, in order to determine the level of difficulty at which to rank each individual tool. I'm not entirely sure how I'm going to do this yet, but right now my standard is simple-- can Google teach me?

Not knowing much about Ngram, I decided to do as all good students do, and immediately check Wikipedia (I'm getting my Masters degree and it hasn't failed me yet, alright?). Here's how Wikipedia describes the tool:
An online search engine that charts frequencies of any set of comma-delimited search strings using a yearly count of n-grams found in sources printed between 1500 and 2008 in Google's text corpora in English, Chinese (simplified), French, German, Hebrew, Italian, Russian, or Spanish; there are also some specialized English corpora, such as American English, British English, English Fiction, and English One Million; the 2009 version of most corpora is also available.
The program can search for a single word or a phrase, including misspellings or gibberish. The n-grams are matched with the text within the selected corpus, optionally using case-sensitive spelling (which compares the exact use of uppercase letters), and, if found in 40 or more books, are then plotted on a graph.
This seems to be quite a powerful tool! What interested me most in this blurb is that the search engine has been programmed to work with such an amazingly large corpus. Although much modern work is still within copyright, it's amazing to be that one program can harness books written over the course of 500 years-- that's astronomical! From this alone, it seems that Ngram will be quite useful to my work.

My next step was to Google "Google Ngram tutorial" and see what I could learn. The first result seemed helpful, so I clicked and this is what I found.

Following the instructions on the webpage, I went to Ngram viewer, typed in a few phrases, and chose a time frame. Because I'm sticking with my dystopian literature theme, I tried to use phrases that I thought would lead to helpful results, and made my time frame span from 1850 to 2000.

The following screen grab shows my results from messing around a bit with the program. It's quite interesting, although I'm surprised that my keywords aren't more successful-- although maybe I'm just not understanding the results. I'm going to tweet out the link to my blog and see if anyone with more knowledge of Ngram responds.

Here are the results from my first searches:


Just from these results, it's interesting to me that the phrase "utopia" spiked in the 1960s, and this is the kind of thing that would lead research questions. In the case of a high school student, this could be a spark that would lead to research for a paper topic. Already, there are good, useful reasons for a teacher to delve into this program.

Lifewire (see above link) also had helpful information for drilling down into more specific tag related searches, which I tried with the search term "Big Brother_NOUN" -- as to differentiate Orwell's all-seeing government from books about familial relations.


How cool is this?? 1984 was published in 1949 and, low and behold, the term spiked around that time period, before dropping off and slowly climbing again. So interesting!

I decided to play around a bit more, and in doing so I found another cool use of the tags feature. I searched "Orwellian" earlier and was less than impressed with the search results. However, this time around I searched "Orwellian_NOUN" and found much more to talk about. Quite interesting how the term has spiked in use in the past 30+ years...


In my browsing, I also found that Google's Ngram help page was useful in picking up some more tips and tricks about the program, such as the following search enhancers:

So much to learn! So much power to harness!
From Google's help page, I learned about the => modifier tag, which tracks term dependencies. For example, on the page the writer explains how the word "tasty" often modifies the word "dessert," there one might search tasty=>dessert. For my purposes I searched utopian=>society:


Strangely, I couldn't find significant results for "dystopian=>society" but the search continues!

Luckily, I had success in my search for "dystopia" as the root of the sentence in this next example, in which I obtained results via the comment _ROOT_=>dystopia. As you can see: in 1994-1995, the topic of dystopia spiked:


Another command recommended by the Ngram info page was [entry]=>*_NOUN, which takes the entry you put in, and uses the * in order to fill in the top ten noun substitutions for the search. For example, I searched utopia=>*_NOUN and my results showed:


As you can see in this graph, the top ranking results is "utopia=>Morris", which spiked in 1977. Who's Morris? I have no idea, but this is the reason that this program is great for research!

Here are some more results for dystopia=>*_NOUN, with the graph settings adjusted slightly:


If you can't tell, I am extremely excited by this adventure into Ngram and am quite hopeful that this will be excellent for my thesis work!

---
*Semi-unrelated aside:

Although my thesis work is based in literature, I was interested in the implications of using Ngram to track the intricacies of language. The very first sentence of the Lifewire link reads:
"A Ngram, also commonly called an N-gram is a statistical analysis of text or speech content to find n (a number) of some sort of item in the text. It could be all sorts of things, like phonemes, prefixes, phrases, or letters."
"Phonemes, prefixes" caught my interest immediately. Outside of literature, I am incredible interested in phonetics and the sounds and pronunciations that make up English, among other languages. Ngram may prove to be useful to me in other areas in my life, it would be cool to see how it could be used to trace language throughout the years. The above-mentioned Google link also had a great deal of information regarding how the program might be used to track language trends. Excuse the aside, but that's what I love about research and learning, you're never done falling down new rabbit holes!

Wednesday, September 13, 2017

NetSmart and the start of a new semester

Source

I can't believe the summer flew by so quickly, but I'm glad to be entering the start of my thesis work, and the beginning of my last year in my Masters degree journey! Hello all, you know me but, if you're reading this and you don't know me, my name is Marissa Candiloro. This blog is going to be dedicated to my first semester of thesis work, which we are calling #ResNetSem, and will be filled with my thesis research progress, along with responses to readings and the events of the semester.

I'll start by talking about my interests and goals in regard to my thesis. First of all, my intention post-graduation is to find a teaching job in a private, classical, Catholic, or Christian high school. I am passionate about teaching English (literature and writing) as well as the atmosphere and mission of such schools. My goal in writing my thesis is to tie my interests into a marketable project that I can show to future employers.

As for my interests, in my time at Kean I have been introduced to a group of fairly new methodologies that are aggregated under the title of "Digital Humanities" or "DH." The DH field includes many different methodologies such as mapping, text mining, and visualizations, to name a few. These methodologies serve as vehicles with which you can examine data in ways that go beyond human capabilities. My favorite example of anything done using DH methodologies is the following chart:

Read about it here
Last semester, I did an independent study that I called Intro to the Digital Humanities, wherein I read,

blogged, and learned about the field and it's methodologies. I even got to attend THATcampDC 2017, which was a great experience. You can read about my independent study here.

I still consider myself very much a newcomer to this field, however I believe that this status puts me in the unique position to be a newcomer speaking to other newcomers-- that is, teachers who have not yet fully incorporated digital methodologies into their classrooms. I'd like my thesis to be an introductory walk-through of 3 (or so) Digital Humanities methodologies that a high school English teachers might utilize in their classrooms, in order to introduce their students to the field, alongside the traditional lessons in close reading and text analysis. I believe that the modern student's work can be enhanced by the DH. To narrow down the scope of potential tools, I am most interested in visualizations and text analysis.

I plan on choosing a handful of books to accompany my walk-through of DH methodologies and to serve as examples throughout the thesis. DH methodologies could be applied to unpack any genre of literature, I could use Shakespeare or Dickens or Austen, however, this is where I would like to tie in another subject I am passionate about: dystopian literature. In addition to my love of 1984 and Brave New World, and my personal interest in unpacking such texts, I think that dystopian novels introduce an interesting lens to my project. Considering how dystopias are often crafted on advanced technology, fear, and control, this might suggest something about how us traditionalist "liberal arts-types" feel about bringing the digital into our text based work. I need to work through my ideas, but I'm excited to see where this idea takes me.

---

Regarding this week's reading, I am so excited to see Howard Rhinegold's work pop up again! I have read some of Net Smart, and I have an immense amount of respect for his work. I think it's fascinating that Rhinegold dove, head first, into the digital world when it was in its infancy, and it's amazing to read his thoughts on how far it has taken us into the future.

I Want Sin: Finding Personhood Amidst Technology in Young Adult Dystopian Literature

I am excited to announce that my thesis, "I Want Sin: Finding Personhood Amidst Technology in Young Adult Dystopian Literature," h...