Anyone who has followed a proposed connection between two passages has faced the obvious objections, that the similarity might be coincidence, or language both writers inherited from a common source or tradition. The best known proposal is a set of seven tests, published in 1989 for assessing echoes of the Hebrew Bible and the Greek Septuagint translation in the letters of Paul and since applied across the New Testament, the Hebrew Bible, and ancient Near Eastern literature. These tests ask whether the source was available, how much language is repeated, how often the same passage is used elsewhere, and how well the echo fits the argument around it. They also ask whether the writer and his audience could plausibly have caught it, whether earlier interpreters heard it, and whether the reading that results makes sense.[1]

Laid out, these rules look like a scorecard, and are frequently used as one. They were not offered that way, and the criteria follow an admission that getting precise results is practically impossible, because interpretation is as much of an art as it is an exact science, and they are offered as rules of thumb instead of a procedure.[2] Three decades of argument about them have mostly been about how much weight any of them can bear.

Heuristics, Not Deduction

The framework is described as similar to listening to music, borrowed from modern literary criticism, and gauged on sensibility more than on method.[3] One such test even comes with a warning from a pioneering literary critic. Asking if earlier interpreters heard the same echo is considered one of the least reliable tests, and should rarely exclude an echo that is considered more likely on other grounds.

One example of a test is to investigate the proposed echo of Job 13:16 in Philippians 1:19 which produces a mixed result, strong on availability and on fit with the argument, moderate on repeated language, and uncompelling on recurrence. Paul plausibly intended it, his audience at Philippi may not have understood it, and no commentator ancient or modern has read it that way. The verdict must be honestly limited to uncertainty. There are always unanswered questions when the criteria are applied to particular texts, and there are definitely moments where the tests fail entirely.[4]

Volume Is the Test Everyone Leans On

Among the criteria, the test for repeated language is more reliable, because shared language can be collected and counted. Extensive overlap makes borrowing easy to recognize, as in Jonah's complaint, which echoes the description of God given to Moses at Sinai.

Exodus 34:6
6The Lord passed by before him and proclaimed: “The Lord, the Lord, the compassionate and gracious God, slow to anger, and abounding in loyal love and faithfulness,
Jonah 4:2
2He prayed to the Lord and said, “Oh, Lord, this is just what I thought would happen when I was in my own country. This is what I tried to prevent by attempting to escape to Tarshish, because I knew that you are a gracious and compassionate God, slow to anger and abounding in mercy, and one who relents concerning threatened judgment.

The comfortable cases can be misleading, because there is no threshold a connection has to clear. The reference in Jonah would be no more allusive if it had also picked up the term for faithfulness from Exodus, and no less allusive if it had dropped an element it does carry, since a marker can be a single word or a name. The theory of allusion behind the criteria requires no particular length or verbal precision.[5]

Resistance to treating repeated language as decisive comes mainly from the study of Mesopotamian and Ugaritic literature, where allusions can be subtle. A demand for substantial overlap is a handy way of ruling out spurious parallels, and the caution is that the principle stops being useful at the point where it is applied rigidly and treated as an unbreakable rule. A Neo-Assyrian oracle asks whose wing has not been clipped, and appears to allude to the story in which Adapa broke the wing of the south wind. Beyond that one image there is hardly any similarity between them.[6]

Availability Assumes a Settled Chronology

The first test asks if the earlier text was available to the later writer, which is straightforward for a New Testament author quoting the Hebrew Bible or Septuagint and difficult anywhere else. The limit is built into the criterion, since analyses of echo are possible only where the relative dating is not controversial.[7] Inside the Hebrew Bible it is rarely possible to establish the priority of one passage over another with confidence, partly because passages that look early often contain later material. Availability is therefore not a prerequisite for identifying a connection, only for working out which passage is referring to which.[8]

That question has criteria of its own, with the same character. A survey of the arguments used to establish a direction of borrowing gathered eight, and the finding was that no one of them applies everywhere or carries the same weight from case to case. Sometimes both passages sit equally well in their contexts, so the test for contextual awkwardness has nothing to measure. Sometimes one criterion supports opposite conclusions about who borrowed from whom.[9]

False Positives and False Negatives

The book of Esther supplies an example of a criterion fighting the very question being asked. Very little published work makes a case for Esther being known and used in Second Temple Judaism, so the availability test returns nothing. The literature is thin because Esther is usually considered to be unknown to early Christianity, which removes a reason to investigate its reception. Applied strictly, the first three tests rule out any proposed echo of Esther, which becomes a circular argument. The recurrence test relies on the same assumption, since the pattern needs a body of work large enough to show it, which much of the New Testament does not offer.[10]

Seven, Two, Eight, or Eleven

Another issue is that there is no single agreed list of tests. In one formulation, only two of the seven tests, availability and volume, are necessary, because the remaining five overlap enough to be ignored.[11] A competing framework for detecting allusions to sayings of Jesus includes eleven items, among them formal agreement and dissimilarity to Graeco-Roman and Jewish traditions, neither of which appears in the original set.[12] A widely cited set of eight guidelines for inner-biblical allusion concerns shared language and nothing else, as the more objective and verifiable evidence.[13]

Even the three tests found in most lists, shared language, shared content, and formal resemblance, get applied very differently. Nearly everyone treats lexical similarity as the soundest, and some as the only real measure. A sizeable group works from content alone, and formal resemblance is rarely used at all.[14] The biggest objection concerns none of the criteria individually, and asks instead where an echo is located, and specifically whether Paul's original audience was able to identify any. Five of the criteria establish that Paul intended the allusion rather than that anyone heard it.[15]

The Same Criteria Produce Different Results

A comparison of three studies of allusion, all claiming precise criteria and all following the same ones, found their results quite different.[16] Competence differs from one interpreter to the next, so one will perceive an allusion, another only a faint echo, and a third no reference at all.

The weighing of evidence, and with it the identification of allusions, has been called an art rather than a science. A rigid mechanical application of any set of criteria is unhelpful, because they are only guidelines and do not hold true in every situation.[17] A person can assemble many lexical, thematic, and formal links between two or more passages and argue forcefully for authorial intent. The next reader may still see a collection of coincidences where the first saw an intentional allusion. What the criteria can only establish is relative probability, whether a proposed allusion falls within the realm of historical and literary possibility.[18]

The Limit of Heuristics

Method of this kind functions as a tentative guide, not a checklist for guaranteed results.[19] A key thing to remember is that every existing set of criteria was assembled out of the passages its author happened to be studying, and no model fits all of scripture. One survey of the options ends on common sense serving better than complex and confusing technology.[20]

General criteria work better as the starting point for a narrower investigation. Deutero-Isaiah reworks Jeremiah by reversal, reprediction, and the reshaping of language, patterns uncommon outside Isaiah 35 and 40 through 66. Those function as criteria, but only inside the collection of texts that produced them.[21]