Skip to content

Concordance, vocabulary & n-gram #101

Description

@sethwoodworth

from: gitberg-temp/issues/20 @whitten

Is it possible to have some tools that would give you various measures on the text.
Such as a concordance of all the words in the archived text with the location/page number,
As well as n-grams for the work like google does, telling you each word (1-gram), each couple of words (2-gram) etc, with the number of times the word shows up in the work.
I'm sure the tools exist, I'm not sure where though.
I do think the resultant files would be a good addition to the repository for each text.

Activity

  1. sethwoodworth commented on Mar 7, 2016

    @sethwoodworth
    ContributorAuthor

    Values that can be expressed in a single line of our metadata.yaml file have a far better chance of getting included into the repo directly. Examples like the Flesch-Kincaid readability score. I like to accept the top 5 keywords found by TF/IDF, but I'd want to discuss that with folks first.

    In general, I think our goal should be to be able to get a book and at least start generating these values from the gitberg tools. At the very least, we can include an Asciidoc > plaintext preprocessor for feeding texts into text analysis engines.

    If a standard format existed for text analysis I would consider generating alongside our ebook generation process, but I'm not aware of any standards for NLP output.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions