Importing from Wikidot
pwikit imports a Wikidot site from a backup made with wikitCLI. The import writes the site's pages with their history, ratings, tags, attachments and forum into a site you have created in pwikit, together with the Wikidot accounts that took part. The original members can then claim their accounts and keep everything they wrote.
What is imported
| Content | What comes across |
|---|---|
| Pages | Name, title, lock state, creation time, time of the last edit, and the account that created the page |
| History | Every revision with its number, author, time and edit comment, and the page source of each revision the backup kept |
| Ratings | Every vote by an account the backup describes, under both Upvote/downvote and Stars |
| Tags | See Tags |
| Parent pages | The parent of each page, when the parent page is also in the backup |
| Attachments | The files, with their names, types, sizes, uploader and upload time |
| Accounts | Every Wikidot account the backup lists, as an unclaimed account |
| Forum | Forum categories; threads with title, description, author, pinned and locked state; posts with their replies and every edit of each post; the discussion thread of each page |
Imported pages can be found by search as soon as the import finishes.
What is not imported
- Site settings such as the title, home page, language and time zone, and themes, roles, permissions and page category settings. Set these in the admin panel.
- The time each vote was cast.
- Page source for revisions the backup kept no source for, such as renames and tag changes. These revisions appear in the history without source.
- The original wikitext of forum posts. The backup stores posts as HTML, which the import converts back to wikitext, so formatting can differ from what was typed.
- Links between pages. Backlinks and
[[module WantedPages]]list links only from pages that have been saved in pwikit since the import. - Votes by accounts the backup does not describe. Revisions, attachments and posts by such accounts are kept without an author.
The backup
Unpack the backup first. pwikit import reads a directory, not an archive file. A wikitCLI backup is one directory per site plus a _users directory holding the accounts; put both into the same directory that the import reads.
A directory is a site when it holds meta/site.json. Both of these work:
One site. Name the site directory itself. Its name is the site's identifier.
xxx-wiki/ the directory given to pwikit import
meta/
site.json the site's own details; this file marks a site directory
pages/*.json one file per page: name, title, parent, tags, rating, lock, revisions, votes, attachments
forum/category/*.json forum categories
forum/<category>/<thread>.json threads
pages/*.7z the page sources, one archive per page, named as in meta/pages
forum/<category>/<thread>.7z the posts of each thread
files/<page name>/<file id> the attachments; a colon in a page name is written %3A
_users/*.json the accounts this backup carries, if it carries anySeveral sites. Unpack every part into one directory and name that directory. Each site keeps the layout above, and -from names which one to import.
wikidot-backup/ the directory given to pwikit import
_users/*.json the accounts, shared by the sites next to it
xxx-wiki/ one site; the directory name is its identifier
xxx-wiki-cn/ another site
xxx-wiki-jp/ another siteA site directory may carry a _users/ of its own as well. Both are read; where the two disagree about an account, the newer fetch wins.
Note The accounts usually sit in a
_users/outside the site directory, next to it. Naming only the site directory still reads the_users/next to it. If the site directory was moved somewhere on its own, with no_users/beside it, those accounts are not found: with no accounts at all the import refuses to run and writes nothing, and with only the ones the site directory carries itself, the other authors are missing. TheN accountsin the summary at the end says how many were created.
The simplest arrangement is to put _users/ and the site directories straight into the archive/ directory in the data directory, which pwikit createsite creates, and run pwikit import without a directory. archive/ then has the several-sites layout above, and a single site needs no -from. Any other directory works too when it is named.
Order of operations
- Install pwikit. See Deployment.
- Create the site with
pwikit createsite. Do not write the starter pages yet. - Take a backup with
pwikit backup create, so that an import you are not satisfied with can be undone. See Operations. - Import the Wikidot backup with
pwikit import. - Create the first administrator with
pwikit admin create, giving your own Wikidot user name. Your imported account becomes the administrator account. See The first administrator. - Complete the checks after importing, then let other people sign up.
Note Import before anyone creates an account under a name from the Wikidot site, including the first administrator. An account created before the import is not linked to the Wikidot account of the same name. The import adds a separate, unclaimed account that receives that person's pages, revisions, votes and posts. To recover, join the two with
pwikit user merge -from wd:<name> -into <name>; see Command line.
If the starter pages were already written with pwikit seed, the import keeps them and skips the backup's pages of the same name, such as nav:side and nav:top.
Running the import
./pwikit import [directory] [options]On Windows, write .\pwikit.exe instead of ./pwikit.
The directory is the unpacked wikitCLI backup: a site directory, or the directory holding several of them. Without a directory, pwikit reads archive/ in the data directory, wherever the command is run from. Options may come before or after the directory.
| Option | Meaning |
|---|---|
-from <slug> | The site inside the backup. Required when the backup holds more than one site |
-site <slug> | The pwikit site to import into. Required when the database holds more than one site. pwikit site list shows the identifiers |
-no-tags | Leave the tags out |
-no-votes | Leave the ratings out |
-no-files | Leave the attachments out |
-no-accounts | Create no accounts, even when the backup holds some. The imported pages, revisions and posts have no authors, and every vote is dropped. A backup that holds no accounts at all needs it to be imported |
-own-users | Read accounts only from the _users/ inside the site directory, not the shared _users/ next to it. Use it when the shared one belongs to another backup and its accounts are not wanted |
-used-users | Create accounts only for the users named by the imported revisions, ratings, attachments and forum. By default every account read is created, including the users of other sites in a shared _users/. With -no-votes or -no-files, users named only by ratings or attachments are not counted |
-user-backfill | On pages the site already has, fill in the authors and ratings that are missing, following the backup; see Filling in authors and ratings. Cannot be combined with -no-accounts |
-update | Bring the pages the site already has up to a newer backup, and add the ratings, attachments, forum threads and posts that are missing; see Updating from a newer backup. The plan is shown and must be confirmed first |
-yes | With -update, go ahead without asking |
-data-dir <dir> | The data directory. Without a directory, the backup is read from its archive/. Attachments are copied into its files/ directory, and pwikit.toml and the bundled PostgreSQL are looked up there. Defaults to PWIKIT_DATA_DIR, then to the directory that holds pwikit |
-database <url> | PostgreSQL connection string. Defaults to DATABASE_URL, then to database in pwikit.toml, then to the bundled PostgreSQL |
Use the same data directory and database as pwikit serve. Otherwise the pages go into another database, or the attachments land where the site does not look for them. The import can run while pwikit serve is running.
Import the only site in archive/ into the only site of the instance:
./pwikit importChoose one site when archive/ holds several:
./pwikit import -from my-wikiImport into one site of an instance that holds several:
./pwikit import -site wikiImport pages and history only:
./pwikit import -no-votes -no-filesImport a backup kept somewhere else:
./pwikit import /srv/wikidot-backupThe import reports its progress and ends with a summary, where N accounts is the number of accounts created:
importing my-wiki into wiki
6210 accounts, 799 pages
200 of 799 pages
400 of 799 pages
600 of 799 pages
100 threads, 312 posts
...
799 pages, 0 already there, 7410 revisions, 108 parents, 159 files, 6210 accounts
9 forum categories, 498 threads, 1830 posts| Message | Cause |
|---|---|
the backup holds 2 sites, name one with -from | The backup holds several sites and -from is missing |
the backup has no site "<slug>" | -from names a site that is not in the backup |
<data directory>/archive does not exist; ... | No directory was given and the data directory has no archive/. Put the backup there, name its directory, or check -data-dir |
no site under "<path>"; ... | Neither the directory nor any directory right below it holds meta/site.json. The backup is usually at the wrong level |
"<path>" is a file, and pwikit reads the unpacked backup; ... | The directory given is an archive file. Unpack it first |
the backup names authors but holds no accounts, so nothing was imported. ... | The backup records authors but no _users/ was read, usually because the site directory was moved somewhere on its own, with no _users/ beside it. Nothing has been written, so put _users/ next to the site directory, or put both in archive/ and run it with no directory, and run it again. To take the content without its authors, pass -no-accounts |
this database holds 2 sites, name one with -site | The instance holds several sites and -site is missing |
no site with slug "<slug>" | -site names a site that does not exist |
this database holds no site, make one with createsite | No site has been created yet |
Tags
The import creates the tags and tag categories the site does not have yet. With -no-tags, pages are imported without tags.
When tags are imported:
- tag names are converted to lower case;
- a tag written as
category:namegoes into the tag categorycategory; - tags that contain a space are skipped.
Decide before the first run. Running the import again does not add tags to pages that are already on the site.
Ratings and attachments
With -no-votes, pages are imported without ratings. With -no-files, attachments are not copied.
Attachments are copied into files/media/ in the data directory. When the backup lists an attachment whose file is missing, the import prints missing attachment <page>/<file> and continues.
Running the import again does not add ratings or attachments to pages that are already on the site.
Forum
The forum is imported only into a site that has no forum section yet. The import creates a forum section named Imported and places the backup's forum categories in it. The discussion thread of a page is attached to that page. A category in which more than half of the threads are page discussions is marked Holds page comments.
When the site already has a forum section, the import prints the site already has a forum, leaving it alone and skips the forum.
Forum categories, threads and posts keep their Wikidot numbers, so old links of the form /forum/c-<number>, /forum/t-<number> and #post-<number> keep working. When a number is already taken by something else on the instance, a new number is used, and the import lists each one at the end:
2 forum numbers from Wikidot were already taken here and got new ones:
t-17071807 -> t-18334710
post-6884529 -> post-9059395The old numbers listed no longer lead to the imported content; change links that use them by hand if needed. Forum categories, threads and posts created on the site afterwards are numbered after the largest imported number.
Accounts
Every account the backup lists is created, including accounts that appear only in other sites of the same backup. The display name is the Wikidot display name. Imported accounts cannot sign in until they are claimed.
A user named by revisions, ratings, attachments or the forum but with no record in the backup's _users/ is usually one who deleted their Wikidot account. The import creates a placeholder account for each such user, shown as wd: followed by the word for "deleted" in the site language and the user's number, and their revisions, ratings and posts belong to it as usual. Placeholder accounts cannot be claimed.
Accounts belong to the instance, not to one site. When a later import names a Wikidot account that the instance already has, the existing account is used and is not changed, whether it has been claimed or not. The one exception is a placeholder: when a later backup carries that user's record, the placeholder takes the real user name and display name.
Running the import again
The import can be run again with the same backup, or with another backup into the same site:
- Pages that already exist on the site are skipped. They receive no revisions, ratings, tags or attachments from the backup; only their parent page is set to the one the backup records.
- Accounts that already exist are reused.
- The forum is skipped when the site already has a forum section.
- With
-user-backfill, pages that already exist get the authors and ratings they are missing from the backup; see Filling in authors and ratings. - With
-update, pages that already exist are brought up to a newer backup; see Updating from a newer backup.
A second run therefore does not repair an import that stopped partway: the page that was being written may lack attachments, and a forum that was being written stays incomplete. If an import stops, restore the backup taken before it and run the import again.
Filling in authors and ratings
Imported pages can lack authors or ratings when:
- They were imported by an older pwikit, which created no account for users with no record in the backup. Their revisions show as the system, and their ratings were dropped.
- The
_users/read at import was incomplete, or-no-accountswas given.
Run the import again with the same backup and -user-backfill to fill them in:
./pwikit import -from <slug> -site <slug> -user-backfill- Empty authors of revisions, pages, attachments, forum threads, posts and post versions are filled in from the backup, and ratings the backup has but the site does not are added.
- Authors and ratings already there are never changed, so the command can be run again safely.
- Forum posts are matched by thread, order and posting time; a post that does not match is skipped.
- Ratings an administrator reset or deleted are added back as well. To fill in authors only, add
-no-votes.
The command ends by printing what it filled in, for example filled in 2053 revision authors, 751 page authors, 682 ratings, ....
Updating from a newer backup
With a newer backup of the same site, -update brings what the site already has up to the backup:
./pwikit import -from <slug> -site <slug> -update- Pages: a page not changed on this site since the import gets the backup's newer revisions, and its title, lock and tags are updated. A page counts as unchanged when its newest revision here has the same number and time as the matching revision in the backup.
- Pages changed on this site: any action that made a new revision, such as an edit, a new title or new tags. These pages are not updated; they are listed before you confirm and again at the end, to be handled by hand.
- New pages: imported in full, as in an ordinary import.
- Ratings: every page that already exists gets the ratings the site does not have. Ratings already there stay as they are. Ratings an administrator reset or deleted are added back too; add
-no-votesto avoid that. - Attachments: attachments the site does not have are added, judged by file name. Attachments already there or deleted are not added again.
- Forum: new forum categories and threads are added. In threads that already exist, the posts the site does not have are added, matched one by one by posting time and author. Posts already there are not changed. What is added keeps its Wikidot numbers as in a first import; see Forum. A forum that was not created by an import is skipped.
- Accounts: as in an ordinary import.
Pages deleted or renamed on Wikidot are not handled, since the backup cannot tell the two apart. A renamed page is imported as a new page under its new name, and the page under the old name stays on the site.
Before anything is written, the command prints how many pages are new, will be updated or are already up to date, lists the pages changed on this site, and asks Continue? [y/N]; any answer other than y cancels. -yes skips the question.
Note Updating a site that people have already started using is not recommended. Pages changed on this site are not updated, only listed; ratings reset or deleted on this site are added back; and in forum threads people have replied to on this site, new posts can only be matched by posting time. The more activity the site has had, the more is left to be handled by hand. Take a backup with
pwikit backup createfirst.
Claiming imported accounts
An imported account holds every page, revision, vote and post of its Wikidot user. Claiming gives the account a password and makes it a normal account; everything it holds stays with it.
On the sign-up page
Members claim their accounts on the site's sign-up page (/-/signup) by entering their Wikidot user name in Username and confirming a code sent to their Wikidot account. The steps are described in Claiming a Wikidot account.
- The code is issued by an external verification service, so the server running pwikit must be able to reach
wikit.unitreaty.orgon the internet. - Names are compared in folded form: lower case, with each run of spaces and punctuation replaced by one hyphen.
Kakushi CTBandkakushi-ctbname the same account. - The claimed account receives the role set as Role after claiming in Site settings.
With a claim link
An administrator can create a link that claims one account without a verification code: in the admin panel, open Users, select Generate claim link, choose the account, and send the link to its owner. The owner sets a password and keeps the Wikidot user name. The link works once and is valid for three days. See Creating accounts and inviting people.
The first administrator
After the import, claim your own account from the command line:
./pwikit admin create -name "Kakushi CTB"pwikit asks for a password without displaying it, sets it, and gives the account every right on every site of the instance:
Password:
took over kakushi-ctb (#12), imported as "kakushi-ctb"See Command line for usage and caveats.
Checks after importing
- Compare the summary at the end of the import with the Wikidot site: the number of pages, attachments and forum threads. Look out for reported missing attachments.
- Set the home page. A new site opens the page
main, while Wikidot sites usually start atstart. In the admin panel, open Site settings and set Home page to the page the Wikidot site used. The backup records it ashome_pageinmeta/site.json. - Spot-check the history, ratings, attachments and tags of several pages, and Users and Forum sections in the admin panel. To add the starter pages the backup does not have, run
./pwikit seed; it writes only pages that do not exist yet.
