Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Pull out all the stops: Textual analysis via punctuation sequences (1901.00519v2)

Published 31 Dec 2018 in cs.CL, cs.LG, and physics.soc-ph

Abstract: Whether enjoying the lucid prose of a favorite author or slogging through some other writer's cumbersome, heavy-set prattle (full of parentheses, em dashes, compound adjectives, and Oxford commas), readers will notice stylistic signatures not only in word choice and grammar, but also in punctuation itself. Indeed, visual sequences of punctuation from different authors produce marvelously different (and visually striking) sequences. Punctuation is a largely overlooked stylistic feature in "stylometry", the quantitative analysis of written text. In this paper, we examine punctuation sequences in a corpus of literary documents and ask the following questions: Are the properties of such sequences a distinctive feature of different authors? Is it possible to distinguish literary genres based on their punctuation sequences? Do the punctuation styles of authors evolve over time? Are we on to something interesting in trying to do stylometry without words, or are we full of sound and fury (signifying nothing)?

Definition Search Book Streamline Icon: https://streamlinehq.com
References (10)
  1. Available at https://github.com/adamjcalhoun/punctuation.
  2. Available at https://medium.com/~@neuroecology/punctuation-in-novels-8f316d542ec4#.brev0b3w1.
  3. Available at https://medium.com/@neuroecology/what-does-punctuation-tell-us-about-republicans-and-democrats-bd46b9f98220.
  4. Class report, Computer Science 229, Stanford University, available at http://cs229.stanford.edu/proj2015/127_report.pdf.
  5. Available at https://www.gutenberg.org.
  6. Available at https://spacy.io.
  7. Available at https://doi.org/10.1016/B0-08-044854-2/04573-9.
  8. Available at http://www.icm2014.org/download/Proceedings_Volume_IV.pdf.
  9. Class report, Computer Science 224, Stanford University, available at https://pdfs.semanticscholar.org/ab0e/be094ec0a44fb0013d640b344d8cfd7adc81.pdf?_ga=2.215953495.1190289256.1578845031-6826891.1578845031.
  10. Available at http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.5.7680.
User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Alexandra N. M. Darmon (1 paper)
  2. Marya Bazzi (11 papers)
  3. Sam D. Howison (15 papers)
  4. Mason A. Porter (210 papers)
Citations (9)

Summary

We haven't generated a summary for this paper yet.