EnglishFrenchSpanish

OnWorks favicon

ids2ngram - Online in the Cloud

Run ids2ngram in OnWorks free hosting provider over Ubuntu Online, Fedora Online, Windows online emulator or MAC OS online emulator

This is the command ids2ngram that can be run in the OnWorks free hosting provider using one of our multiple free online workstations such as Ubuntu Online, Fedora Online, Windows online emulator or MAC OS online emulator

PROGRAM:

NAME


ids2ngram - generate n-gram data file from ids file

SYNOPSIS


ids2ngram [option]... ids_file...

DESCRIPTION


ids2ngram generates idngram file, which is a sorted [id1,..,idN,freq] array, from binary
id stream files. Here, the id stream files are always generated by mmseg or slmseg.
Basically, it finds all occurrence of n-words tuples (i.e. the tuple of (id1,..,idN)), and
sorts these tuples by the lexicographic order of the ids make up the tuples, then write
them to specified output file.

INPUT


The input file is presented as a binary id stream, which looks like:
[id0,...,idX]

OPTIONS


All the following options are mandatory.

-n,--NMax N
Generates N-gram result. ids2ngram does only support uni-gram, bi-gram, and trigram,
so any number not in the range of 1..3 is not valid.

-s,--swap swap-file
Specify the temporary intermediate file.

-o, --out output-file
Specify the result idngram file, e.g. the array of [id1, ..., idN, freq]

-p, --para N
Specify the maximum n-gram items per paragraph. ids2ngram writes to the temporary file
on a per-paragraph basis. Every time it writes a paragraph out, it frees the
corresponding memory allocated for it. When your computer system permits, a higher N
is suggested. This can speed up the processing speed because of less I/O.

EXAMPLE


Following example will use three input idstream file idsfile[1,2,3] to generate the
idngram file all.id3gram. Each para (internal map size or hash size) would be 1024000,
using swap file for temp result. All temp para result would eventually be merged to got
the final result.

ids2ngram -n 3 -s /tmp/swap -o all.id3gram -p 1024000 idsfile1 idsfile2 idsfile3

Use ids2ngram online using onworks.net services


Free Servers & Workstations

Download Windows & Linux apps

  • 1
    PHP QR Code
    PHP QR Code
    PHP QR Code is open source (LGPL)
    library for generating QR Code,
    2-dimensional barcode. Based on
    libqrencode C library, provides API for
    creating QR Code barc...
    Download PHP QR Code
  • 2
    Freeciv
    Freeciv
    Freeciv is a free turn-based
    multiplayer strategy game, in which each
    player becomes the leader of a
    civilization, fighting to obtain the
    ultimate goal: to bec...
    Download Freeciv
  • 3
    Cuckoo Sandbox
    Cuckoo Sandbox
    Cuckoo Sandbox uses components to
    monitor the behavior of malware in a
    Sandbox environment; isolated from the
    rest of the system. It offers automated
    analysis o...
    Download Cuckoo Sandbox
  • 4
    LMS-YouTube
    LMS-YouTube
    Play YouTube video on LMS (porting of
    Triode's to YouTbe API v3) This is
    an application that can also be fetched
    from
    https://sourceforge.net/projects/lms-y...
    Download LMS-YouTube
  • 5
    Windows Presentation Foundation
    Windows Presentation Foundation
    Windows Presentation Foundation (WPF)
    is a UI framework for building Windows
    desktop applications. WPF supports a
    broad set of application development
    features...
    Download Windows Presentation Foundation
  • 6
    SportMusik
    SportMusik
    Mit dem Programm kann man schnell und
    einfach Pausen bei Sportveranstaltungen
    mit Musik �berbr�cken. Hierf�r haben sie
    die M�glichkeit, folgende Wiedergabvaria...
    Download SportMusik
  • More »

Linux commands

Ad